SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
ML Model and Foundation Model (FM) Development

Training, Hyperparameter Tuning and Fine-Tuning Models on AWS (MLA-C02)

19 min readMLA-C02 · ML Model and Foundation Model (FM) DevelopmentUpdated

Training and customizing models on AWS means turning data into a working model with Amazon SageMaker AI (built-in algorithms or your own script in script mode), tuning its hyperparameters with automatic model tuning, and adapting foundation models with prompt engineering, fine-tuning and better retrieval. Task 2.2 of the MLA-C02 exam tests the decisions inside that work: which algorithm or estimator to use, how to make training faster, how to diagnose overfitting, underfitting and catastrophic forgetting, when to combine models, and which customization method suits the data you actually have. This lesson walks through each in the order you would meet them on a project.

What you’ll learn
  • Choose the SageMaker AI built-in algorithm for a problem type and prepare its data and input channels (format, input mode, distribution).
  • Run existing framework code in SageMaker AI script mode, including hyperparameters, environment variables, model output and extra packages.
  • Configure automatic model tuning: strategy, parallelism, range scaling, custom objective metrics and warm start.
  • Reduce training time with early stopping and the right kind of distributed training.
  • Diagnose overfitting, underfitting and catastrophic forgetting and apply the matching fix, including the effect of epochs, steps, batch size and learning rate.
  • Combine models with bagging, boosting, stacking, cascades or Amazon Bedrock intelligent prompt routing for accuracy or cost.
  • Pick a foundation model customization method (prompt engineering, fine-tuning, continued pre-training, distillation, LoRA) and tune embeddings, chunking and search type for retrieval.

Which SageMaker AI built-in algorithm fits which problem?

Pick a built-in algorithm by the shape of the problem and whether you have labels: each built-in is tuned for one problem type, and the wrong family either needs labels you do not have or ignores structure (time, pairs, sparsity) that matters. Built-in algorithms run in AWS-managed containers, emit their metrics automatically and need no training code, which is why they are the quickest route when one fits.

ProblemBuilt-in algorithmWhy it fits
Tabular classification or regressionXGBoost (also Linear Learner, LightGBM, CatBoost)Gradient-boosted trees are a strong default on tabular data; Linear Learner suits linear relationships and very large sparse data
Anomaly scores on unlabeled recordsRandom Cut Forest (RCF)Unsupervised; learns normal data and returns an anomaly score per record
Unusual entity-to-IP-address pairsIP InsightsUnsupervised, but specific to entity/IP associations, not general sensor or numeric data
Forecasting many related time seriesDeepAROne recurrent model trained across all series, so patterns transfer to new items with short histories
Grouping unlabeled recordsk-meansClusters records; it does not output an anomaly score
Recommendations on sparse user-item dataFactorization MachinesModels pairwise interactions in high-dimensional sparse data
Text classification or word vectorsBlazingTextFast Word2Vec embeddings and supervised text classification
Classification or regression by similarityk-nearest neighbors (k-NN)Predicts from the closest stored examples; no temporal structure

Reading the scenario for two words settles most choices: labeled or not (supervised XGBoost/Linear Learner versus unsupervised RCF/k-means) and what the output must be (a class, a number, a score, a cluster, a forecast). When no built-in fits, or the team already has code in a common library such as scikit-learn, PyTorch, TensorFlow or Hugging Face Transformers, use script mode instead (below). Choosing between built-ins, managed AI services and foundation models in the first place belongs to Task 2.1.

How do you feed training data to a SageMaker AI training job?

A training job reads each input channel from Amazon S3 in one of three input modes and with one of two distribution types, and built-in algorithms also expect a specific file format. Getting the format wrong makes a job fail or train on garbage; getting the mode and distribution wrong makes it slow to start.

Format. Built-in algorithms that accept CSV (XGBoost, Linear Learner and others) expect the target in the first column and no header row, with content type text/csv. A header row is read as data and a label in the last column is treated as a feature. Declaring a different content type (for example application/x-recordio-protobuf) does not convert the file; the declared type must match the real format.

SettingOptionBehaviour
Input modeFile (default)Downloads the whole channel to the instance's storage volume before training starts; large datasets mean a long wait
Input modeFastFileExposes S3 objects as files and streams them on demand, so training starts almost immediately; best for file-based, sequential reads
Input modePipeStreams data through a named pipe; fast, but the script must read a stream rather than files
S3 data distributionFullyReplicated (default)Every instance receives the full dataset
S3 data distributionShardedByS3KeyEach instance receives a separate subset of the S3 objects

Mode and distribution solve different problems and are often combined: FastFile removes the up-front copy, ShardedByS3Key stops every instance from receiving every file. Sharding is only safe when each instance should train on its own portion and the script trains on whatever appears in its channel; a script that already picks its own slice of the full file list would shard twice. Only the input mode changes how long the up-front copy takes; storage size and pricing options leave it as it is.

What is SageMaker AI script mode and how does a script plug into it?

Script mode runs your own training script, unchanged or nearly so, inside an AWS-managed framework container: you choose a framework estimator (PyTorch, TensorFlow, Hugging Face, scikit-learn, XGBoost), point entry_point at the script and set a supported framework_version. AWS builds and patches the container and provisions the managed training instances, so you get custom code without maintaining images. Building your own image and using the generic Estimator (bring your own container) is the alternative when no managed container fits; a notebook instance is an interactive server, not managed training. These are the SageMaker Python SDK v2 names; SDK v3 replaces the estimator classes with ModelTrainer plus a SourceCode configuration (source directory, entry script, requirements file), while the container conventions below stay the same.

The container and the script communicate through fixed conventions:

  • Hyperparameters set on the estimator (hyperparameters={'max_depth': 6, 'dropout': 0.3}) arrive as command-line arguments (--max_depth 6 --dropout 0.3), so the script parses them with argparse. Automatic model tuning passes each trial's values the same way, which is why hard-coded values defeat tuning.
  • Input data: each channel is mounted under /opt/ml/input/data/<channel>, exposed as SM_CHANNEL_<NAME> (for example SM_CHANNEL_TRAINING).
  • Model output: save the trained model to SM_MODEL_DIR, which is /opt/ml/model. SageMaker AI packages that directory as model.tar.gz and uploads it to the estimator's output_path. A model saved anywhere else is lost, and the job still reports Completed with an empty archive.
  • Checkpoints: /opt/ml/checkpoints synced to checkpoint_s3_uri is for resuming (for example after a Spot interruption), not for the final model.
  • Extra packages: put a requirements.txt in source_dir (the folder that holds the entry point). The framework container pip-installs it before running the script, so you keep the managed image. This is the lightest way to add a dependency: the image stays AWS-managed and patched, and nothing else in the job changes.

How does SageMaker AI automatic model tuning choose hyperparameters?

Automatic model tuning (AMT) runs many training jobs over ranges you define and keeps the one with the best objective metric; the strategy decides how each next combination is picked. Choose the strategy from the budget and from whether results should guide the search.

StrategyHow it picks combinationsUse when
Bayesian (default)Treats tuning as a regression problem and uses completed jobs to choose the next values, balancing exploration and exploitationLimited job budget, continuous ranges, results should inform the search
RandomSamples each combination independently of earlier resultsMany jobs can run fully in parallel
GridTries every combination of listed values; categorical parameters onlySmall, discrete search spaces that must be covered exhaustively
HyperbandUses intermediate (per-epoch) metrics to stop weak jobs early and give more resources to promising onesLong iterative training where poor configurations are visible after a few epochs

Parallelism matters for Bayesian search. MaxParallelTrainingJobs sets how many jobs run at once. If all jobs launch together, none can learn from another and Bayesian search degrades to random sampling. With a fixed total (MaxNumberOfTrainingJobs), lower parallelism usually finds a better result, at the cost of wall-clock time.

Ranges and scaling. Ranges are continuous, integer or categorical. The scaling type controls how values are sampled: Linear spreads them evenly, so in a range like 0.0001 to 1 about nine samples in ten land above 0.1 and the small values are barely explored; Logarithmic samples each order of magnitude equally and suits learning rates and regularization strengths (positive values spanning several powers of ten); ReverseLogarithmic concentrates samples near 1 and suits values such as momentum in the range 0 to just under 1.

Objective metric. Built-in algorithms publish their metrics automatically. A script-mode job must print the metric to its logs in a consistent format (for example test_rmse: 4.21), and the estimator or tuning job needs a metric definition with a regex that captures it; you then name that metric as the objective with its direction (Maximize or Minimize). The tuner reads only what that regex finds in the job's logs.

Warm start reuses earlier tuning jobs (parents) instead of starting from scratch. IdenticalDataAndAlgorithm requires the same data and training image, though ranges can change. TransferLearning allows different data, changed ranges and a different algorithm version, so it is the type for a dataset that has grown since the parent job.

How do early stopping and distributed training cut training time?

Training time falls in two ways: stop work that will not pay off (early stopping) and spread the work that will across more hardware (distributed training). Faster input modes (FastFile, sharding) help the start-up phase, covered above.

Early stopping in a tuning job. Setting TrainingJobEarlyStoppingType to Auto applies a specific rule: after each epoch, AMT compares the current job's objective metric with the median of the running averages of the earlier jobs' objectives up to that same epoch, and stops the job if it is worse than that median. That is aggressive: roughly half of the weaker jobs can be cut early. It needs no code change for built-in algorithms that emit per-epoch metrics; a custom script (for example PyTorch) must log the objective metric after every epoch for Auto to work. Caps on runtime or job count are blunt instruments by comparison, because they stop good and bad jobs alike. Hyperband has its own built-in early stopping, so AMT requires TrainingJobEarlyStoppingType to be Off when you choose the Hyperband strategy.

Early stopping inside one job. Many algorithms stop themselves when a validation metric stops improving. For built-in XGBoost, provide a validation channel and set early_stopping_rounds: training ends after that many rounds without improvement, so the stopping point follows the data instead of a round count guessed in advance. Without a validation channel there is nothing to watch, and early stopping cannot trigger.

Distributed training. Choose by where the bottleneck is.

SituationApproachSageMaker AI library
Model fits on one GPU; the dataset is large and training is slowData parallelism: a full model copy on every GPU, each processing different batches, gradients averagedDistributed data parallelism library (SMDDP)
Parameters, gradients and optimizer states do not fit on one GPU, even at batch size 1Model parallelism / sharded data parallelism: the model state is split across GPUsModel parallelism library (SMP v2, built on PyTorch FSDP), including sharded data parallelism and tensor parallelism

Data parallelism does not help an out-of-memory error caused by model state, because it replicates that state on every GPU; techniques that shrink the per-step batch only reduce activation memory, not the memory taken by the model state itself. Conversely, splitting a model that already fits adds communication overhead for no benefit. A bigger single instance is a scale-up, not distributed training. Managed Spot training with checkpoints lowers cost, not time.

How do you prevent overfitting, underfitting and catastrophic forgetting?

Diagnose first from the gap between training and validation metrics, then apply the fix that matches: overfitting needs less capacity or more data, underfitting needs more capacity or better signal, and catastrophic forgetting needs fine-tuning that disturbs the base model less.

ProblemSymptomFixesMakes it worse
Overfitting (high variance)Training metric excellent, validation clearly worse, on data from the same distributionStronger L1/L2 regularization (XGBoost alpha/lambda), row and column subsampling, shallower trees, dropout, early stopping, more or augmented data, fewer featuresDeeper trees, more rounds or epochs, removing the validation set
Underfitting (high bias)Training and validation both poor and close together; learning plateaus earlyMore capacity (deeper trees, more layers), less regularization, more informative features, longer training if still improvingMore subsampling, more regularization, stopping earlier, more rows of the same weak features
Catastrophic forgettingA fine-tuned model is good at the new narrow task but has lost general abilities it used to haveParameter-efficient fine-tuning (LoRA keeps base weights frozen), mixing general instruction data into the training set, lower learning rate, fewer epochsMore epochs or a higher learning rate on the narrow data; more domain-only training

Catastrophic forgetting is overfitting's cousin for foundation models: full fine-tuning on a small, narrow dataset for several epochs overwrites weights that encoded general skills. Debugging a model that will not converge at all, and the evaluation metrics themselves, belong to Task 2.3.

How do epochs, steps, batch size and learning rate interact?

An epoch is one full pass over the training data; a step (iteration) is one optimizer update on one batch; the batch size is how many examples go into each update. So steps per epoch = records ÷ batch size, and total steps = steps per epoch × epochs. With 120,000 records, a batch of 400 and 10 epochs, that is 300 steps per epoch and 3,000 steps in total, the numbers a learning-rate scheduler or evaluation interval needs.

  • Batch size and memory. GPU memory per step grows with the per-device batch size. To keep a large effective batch on limited memory, use gradient accumulation: run several small batches and update once (16 batches of 64 behave like one batch of 1,024). Memory per step is set by the model and the per-device batch; settings that change how long you train or how data reaches the instance do not change it.
  • Batch size and the learning rate. In data-parallel training the global batch is the per-GPU batch times the number of GPUs, so each epoch has far fewer updates. Leaving the single-GPU learning rate in place usually costs accuracy. A widely used heuristic (not an AWS feature) is to scale the learning rate roughly in proportion to the global batch and add a warmup that ramps it up over the first part of training so the larger rate does not destabilize early updates.
  • Epochs. Too few underfit, too many overfit (and, for foundation models, encourage forgetting); early stopping picks the point from validation data rather than a guess.

How do you combine models to improve accuracy or reduce cost?

Combine models for accuracy with an ensemble that fixes the specific weakness, and combine them for cost with a cascade or router that sends easy requests to a cheap model and only hard ones to an expensive model.

TechniqueHow it worksMainly reduces
Bagging (for example random forest)Many copies of the same model type trained in parallel on bootstrap samples; outputs averaged or votedVariance: unstable models that swing between retrains
Boosting (for example XGBoost)Weak models (shallow trees) added in sequence, each fitted to what the ensemble still gets wrongBias: models too simple to fit the data
StackingA meta-model trained on the out-of-fold predictions of several different models learns how much to trust eachErrors of diverse models that are each strong in different regions
Voting / averagingFixed rule (majority or fixed weights) over several modelsSome error, but cannot learn when to trust which model

Cascades for cost. When a small, cheap model is reliable on most traffic and its confidence scores are well calibrated, score everything with it first and escalate only low-confidence cases to the large model. The expensive model then sees a fraction of traffic, while hard cases keep its accuracy. The saving depends on routing by difficulty: any scheme that sends every request to both models, or splits traffic without looking at difficulty, keeps either the full cost or the small model's errors.

Amazon Bedrock intelligent prompt routing is the managed version for foundation models: a prompt router sends each request to one of several models in the same model family, predicting which can answer well at lower cost, so you do not build or maintain your own request classifier. Capacity options change how much traffic one model can serve, not which model answers a request.

How do you customize a foundation model: prompt engineering or fine-tuning?

Start with task-specific prompt engineering and move to training only when prompting cannot deliver or costs too much. Prompting changes the input; fine-tuning, continued pre-training and distillation change the model.

Prompt engineering needs no data and no training job: clear instructions, a system prompt for role and rules, the exact output format or JSON schema, and a few input/output examples (few-shot). It is the fastest, cheapest fix when the model already knows the task but behaves inconsistently, for example when the format drifts between calls. Sampling parameters can make output less random, but they cannot specify a structure the prompt never asked for.

Method (Amazon Bedrock)Training dataUse when
Fine-tuningLabeled prompt-completion pairsA stable task, style or format with many good examples, especially when few-shot prompts have grown long and expensive
Continued pre-trainingUnlabeled, input-only textThe model must learn domain vocabulary and language it misreads
Model DistillationYour prompts; the teacher model generates the responsesA smaller, cheaper, faster student model in the same family should perform close to a large teacher on your task, and you have prompts but no labeled responses
Custom Model ImportWeights trained elsewhereHosting a model you customized outside Bedrock; it trains nothing

Fine-tuning moves repeated examples out of every prompt and into the model, so calls become short. A custom Bedrock model needs an inference option: on-demand deployment where the model supports it, otherwise Provisioned Throughput, whose fixed cost pays off only at sustained volume. RAG is for knowledge that changes, not for a fixed style.

Parameter-efficient fine-tuning (PEFT) for open-weight models in SageMaker AI (for example from SageMaker JumpStart): LoRA freezes the base weights and trains small low-rank adapter matrices, which removes most gradient and optimizer memory, limits catastrophic forgetting and produces a small artifact per task instead of a full model copy. Loading the frozen base in quantized form while training adapters cuts memory further so the job fits on a smaller GPU: QLoRA specifically quantizes the base to 4-bit (NF4), and 8-bit loading with LoRA adapters is a related option. The memory saving comes from what is trained and how the base is stored, not from how data is spread across GPUs.

How do you optimize retrieval components and embedding models?

Retrieval quality in a RAG system depends on three tunable parts: the embedding model that turns text into vectors, the chunking that decides what each vector represents, and the search type that matches queries to chunks. Diagnose which one causes the failure before changing anything; choosing RAG as the architecture is Task 2.1, and reranking and knowledge base setup are covered in Domain 3.

Embedding models. Amazon Titan Text Embeddings V2 lets you choose 256, 512 or 1,024 output dimensions (and whether to normalize vectors). Fewer dimensions mean smaller vectors, a smaller index and faster search for a small loss of accuracy. When a domain's terms are confused (different compound or product names landing close together), fine-tune the embedding model on query and relevant-passage pairs with a contrastive loss; only the embedding model decides where texts land in the vector space, so changes elsewhere in the pipeline cannot pull confused terms apart. Every vector in an index must come from the same model and dimension: after changing either, re-embed the whole corpus, re-embed queries with the same model, and rebuild the index. Mixing old and new vectors breaks similarity.

SymptomCauseChange
Exact identifiers (part numbers, codes, names) return similar but wrong chunksVector similarity alone blurs exact tokensHybrid search (keyword plus semantic) instead of semantic-only
The right chunk is found but lacks surrounding contextChunks too narrowHierarchical chunking: search small child chunks, return their larger parent chunk
Key sentences are cut in half at chunk boundariesNo overlap between fixed-size chunksAdd overlap (a percentage of each chunk shared with its neighbour) through a data source configured with it, then ingest again
Chunks mix unrelated topicsBoundaries ignore meaningSemantic chunking, which splits where the meaning changes
Storage cost and latency too high, small accuracy loss acceptableLarge vectorsSmaller embedding dimension, re-embed and rebuild

In Amazon Bedrock Knowledge Bases, chunking (default, fixed-size with overlap, hierarchical, semantic, none, or a custom Lambda transformation) is fixed when the data source is created and cannot be edited afterwards, so changing it means creating a new data source (or knowledge base) with the new configuration and ingesting the documents again; the search type (hybrid or semantic) can be overridden per query on supported vector stores such as Amazon OpenSearch Serverless. Very small chunks create more boundaries and lose context, and returning more results adds noise without fixing either cause.

Tip. Task 2.2 questions are scenarios that describe a training job, tuning job, fine-tuning run or RAG pipeline that is slow, failing, inaccurate or too expensive, often with one constraint in the middle of the story: no labels, no custom container images, a fixed number of training jobs, no larger instance, no classifier to maintain, or an acceptable accuracy loss. Expect near-twin options that differ in one setting (Logarithmic versus ReverseLogarithmic scaling, TransferLearning versus IdenticalDataAndAlgorithm, FastFile versus a bigger volume, hybrid versus semantic search, data versus model parallelism), and multi-response items where two settings together solve two separate symptoms. Read the training and validation numbers carefully before deciding between overfitting and underfitting, and check what data the scenario actually has before choosing a customization method.

Key takeaways
  • Built-in algorithm choice follows labels and output: RCF for unlabeled anomaly scores, DeepAR for many related time series, XGBoost for tabular prediction, k-means only clusters.
  • Built-in CSV input means the label in the first column and no header; FastFile streams data and ShardedByS3Key gives each instance its own subset.
  • In script mode, hyperparameters arrive as command-line arguments, the model must be saved to SM_MODEL_DIR (/opt/ml/model), and requirements.txt in source_dir adds packages to the managed container.
  • Bayesian tuning learns from completed jobs, so high parallelism turns it into random search; use logarithmic scaling for ranges spanning orders of magnitude and a regex metric definition for custom scripts.
  • Warm start TransferLearning handles new data or a new algorithm version; IdenticalDataAndAlgorithm needs the same data and image.
  • Early stopping (AMT Auto, Hyperband, XGBoost early_stopping_rounds with a validation channel) saves wasted compute; use data parallelism when the model fits on one GPU and model parallelism or sharding when it does not.
  • Overfitting needs regularization or data, underfitting needs capacity or features, and catastrophic forgetting needs LoRA, mixed general data, a lower learning rate or fewer epochs.
  • Bagging reduces variance, boosting reduces bias, stacking learns how to weight diverse models, and cascades or Bedrock intelligent prompt routing cut cost by sending easy requests to smaller models.
  • Prompt engineering first; fine-tuning needs labeled pairs, continued pre-training uses unlabeled text, and distillation trains a smaller student from a teacher's responses to your prompts.
  • Changing the embedding model or dimension means re-embedding the corpus; hybrid search fixes exact identifiers, hierarchical chunking adds context, and overlap stops split sentences.

Frequently asked questions

What is SageMaker AI script mode?

SageMaker AI script mode runs your own training script inside an AWS-managed framework container such as PyTorch, TensorFlow, Hugging Face or scikit-learn. You pass the script as entry_point to the framework estimator with a supported framework_version; SageMaker AI provisions the instances, passes hyperparameters as command-line arguments, mounts input channels under /opt/ml/input/data and uploads whatever the script saves to /opt/ml/model as model.tar.gz. Extra Python packages go in a requirements.txt file in source_dir, so you never build or patch your own image.

Which automatic model tuning strategy should I use in SageMaker AI?

Use Bayesian optimization when the job budget is limited and each new combination should learn from completed jobs; keep parallel jobs low so it can. Use random search when many jobs will run fully in parallel, grid search only for small sets of categorical values, and Hyperband for long iterative training where poor configurations are visible after a few epochs, because it stops weak jobs early and gives more resources to promising ones.

How do I know whether a model is overfitting or underfitting?

Compare training and validation metrics on data from the same distribution. If training performance is excellent but validation is clearly worse, the model is overfitting and needs regularization, subsampling, dropout, early stopping or more data. If both are poor and close together, the model is underfitting and needs more capacity, less regularization or more informative features. Adding more rows of the same weak features does not fix underfitting.

What is catastrophic forgetting in fine-tuning?

Catastrophic forgetting is when a foundation model fine-tuned on narrow data becomes good at the new task but loses general abilities it had before, such as summarization or following ordinary instructions. It is most likely with full fine-tuning on a small dataset for several epochs. Mitigations are parameter-efficient fine-tuning such as LoRA (base weights stay frozen), mixing general instruction data into the training set, a lower learning rate and fewer epochs.

What is the difference between fine-tuning and continued pre-training in Amazon Bedrock?

Fine-tuning in Amazon Bedrock trains a model on labeled prompt-completion pairs to perform a specific task, style or format. Continued pre-training trains it on unlabeled, input-only text so it learns a domain's vocabulary and language. If you only have raw documents, continued pre-training is the option; if you have examples of the inputs and the outputs you want, fine-tune. Model Distillation is a third option that uses a large teacher model to generate the responses a smaller student learns from.

How do I calculate steps per epoch?

Steps per epoch equal the number of training records divided by the batch size, and total steps equal steps per epoch multiplied by the number of epochs. For example, 120,000 records with a batch size of 400 gives 300 steps per epoch, and 10 epochs gives 3,000 steps. In data-parallel training, use the global batch size (per-device batch size times the number of devices), which is why each epoch has fewer steps when you add GPUs.

When should I use data parallelism versus model parallelism on SageMaker AI?

Use data parallelism, with the SageMaker AI distributed data parallelism library, when the model fits on one GPU and training is slow because the dataset is large: each GPU holds a full model copy and processes different batches. Use the SageMaker AI model parallelism library, including sharded data parallelism, when parameters, gradients and optimizer states do not fit on one GPU even at a batch size of 1, because it splits that state across devices.

Do I need to re-embed my documents if I change the embedding model?

Yes. Vector similarity only works when queries and documents are embedded by the same model with the same output dimension, so changing the embedding model, fine-tuning it or changing the dimension (for example Amazon Titan Text Embeddings V2 from 1,024 to 512) means re-embedding the whole corpus and rebuilding the vector index. Mixing old and new vectors in one index makes similarity scores meaningless.

Source

This lesson covers the "ML Model and Foundation Model (FM) Development" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.

Test yourself on this topic
Practice questions with full explanations.
Practice now

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?