SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
MLA-C02 · Domain 2

ML Model and Foundation Model (FM) Development practice questions

ML Model and Foundation Model (FM) Development is worth 24% of the MLA-C02 exam — the 2nd-heaviest of the 4 domains. Choosing modeling approaches — traditional ML, managed AI services, foundation models and RAG — then training, fine-tuning and customizing models, and evaluating ML and GenAI performance. 6 fully worked examples are further down this page, answers included.

Exam weight
24%
the 2nd-heaviest of the 4 domains
Questions
114
across 3 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample ML Model and Foundation Model (FM) Development questions, fully explained

Questions from the MLA-C02 bank mapped to domain 2, with the answer key and the reasoning behind every option. None of them repeat the examples on the main MLA-C02 practice page.

Question 1ML Model and Foundation Model (FM) Development

An HR team wants an internal assistant that answers employee questions about company policies. Policies are revised most weeks, and every answer has to show which policy document it came from so employees can check the source. The team wants to avoid retraining a model whenever a policy changes. Which approach should the ML engineer choose?

Choose one.

  • a
    Fine-tune a foundation model on the policy documents each week

    Weekly fine-tuning is the retraining the team wants to avoid, and a fine-tuned model does not return the source document for its answers.

  • b
    Run continued pre-training on the policy corpus each quarter

    Quarterly training would answer from policies that are weeks out of date, and continued pre-training produces no citations.

  • c
    Paste summaries of the current policies into a fixed system prompt

    Summaries drift from the real documents and must be edited by hand after each revision; a fixed prompt also gives no per-answer source document.

  • d
    Connect the policy documents to an Amazon Bedrock knowledge base Correct

    Retrieval Augmented Generation fetches the current policy text at question time and returns citations; re-syncing the data source picks up revisions without retraining a model.

The concept

RAG grounds a model in documents retrieved at query time, which suits frequently changing knowledge and answers that need citations.

Why that’s the answer

Two requirements decide it: knowledge that changes weekly and answers that cite their source. A Bedrock knowledge base retrieves the relevant policy chunks for each question and returns citations, and a data source sync keeps it current with no model training. Fine-tuning and continued pre-training both bake a snapshot of the policies into the weights, go stale between runs and cannot point to a source. Hand-maintained prompt summaries are fragile and uncited.

How to reason it out
  1. Note that the facts change frequently.
  2. Note that answers need a traceable source.
  3. Recall that RAG retrieves current documents at query time and returns citations.
  4. Reject options that store knowledge in model weights.

Exam tip: Changing facts plus citations points to RAG, not fine-tuning.

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02) — the lesson that teaches this.

Question 2ML Model and Foundation Model (FM) Development

A travel company uses a large foundation model in Amazon Bedrock to extract booking details from customer emails. Accuracy is good, but latency and cost are too high at current volume. The team has 8,000 production prompts in its invocation logs, but no human-written reference answers and no automated way to score a response. Which customization approach fits these constraints?

Choose one.

  • a
    Use Amazon Bedrock Model Distillation with the large model as the teacher Correct

    Distillation generates responses from the teacher for the team's prompts and fine-tunes a smaller student on them, giving a faster, cheaper model for this task without hand-written labels.

  • b
    Run supervised fine-tuning of a smaller model on the logged prompts

    Supervised fine-tuning needs labeled prompt-response pairs; the team has prompts but no reference answers to train toward.

  • c
    Run reinforcement fine-tuning of a smaller model with a reward function

    Reinforcement fine-tuning learns from reward scores; the scenario states there is no automated way to score a response, so there is nothing to write the reward function from.

  • d
    Purchase Amazon Bedrock Provisioned Throughput for the large model

    Provisioned Throughput reserves capacity for the same large model; it does not make each request faster or cheaper to compute.

The concept

Model distillation transfers a large teacher model's behaviour on a specific task to a smaller, faster student model, using the teacher's own responses as training data.

Why that’s the answer

The constraints are prompts without reference answers and no scoring logic. Bedrock Model Distillation needs exactly what the team has: use-case prompts (or invocation logs) plus a teacher whose accuracy is acceptable. It generates teacher responses and fine-tunes the student on them. Supervised fine-tuning needs labeled pairs, reinforcement fine-tuning needs a reward function, and Provisioned Throughput keeps the same expensive model.

How to reason it out
  1. Identify the goal: the large model's accuracy at lower latency and cost.
  2. Inventory the data: prompts only, no labels, no grader.
  3. Match the data to the method: distillation uses teacher responses as labels.
  4. Reject methods that need labels or a reward function.

Exam tip: Good teacher, prompts but no labels, need smaller and faster: distill.

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02) — the lesson that teaches this.

Question 3ML Model and Foundation Model (FM) Development

A manufacturer asks a foundation model in Amazon Bedrock to write programs in its in-house machine-control language. The language reference is already supplied through a knowledge base, yet most generated programs fail the compiler or the unit-test suite. No team has written reference solutions, but a test harness can compile any program and return a pass rate. Which customization approach should the ML engineer use?

Choose one.

  • a
    Supervised fine-tuning on prompt and completion pairs cut from the language reference manual

    Manual excerpts teach the same reference text that retrieval already supplies; they are not examples of correct programs for real tasks, which is what supervised fine-tuning would need.

  • b
    Model distillation that uses a larger model from the same family as the teacher model

    A student learns to copy its teacher. The larger model has never seen this in-house language either, so distillation transfers its failing programs.

  • c
    A reranking model that reorders the knowledge base passages before each generation

    Retrieval already supplies the reference, and the failures are in the generated code. Better-ranked passages do not teach the model to produce programs that pass the tests.

  • d
    Reinforcement fine-tuning with an AWS Lambda reward function that grades each generated output Correct

    Reinforcement fine-tuning learns from reward scores instead of labeled pairs; a Lambda reward function that runs the harness turns its pass rate into the training signal.

The concept

Reinforcement fine-tuning improves a model from feedback produced by a reward function, so it fits tasks where outputs can be scored automatically but correct answers were never written down.

Why that’s the answer

There are no labeled reference programs, which rules out supervised fine-tuning on real examples, but an automatic grader exists. Reinforcement fine-tuning in Amazon Bedrock takes training prompts and a reward function (for example in AWS Lambda), so the compile-and-test pass rate becomes the learning signal. Supervised fine-tuning is the near-twin, but manual excerpts are not examples of correct programs. Distillation is tempting because it needs no labels, but a student can only be as good as its teacher, and no available teacher knows the language. Better retrieval ranking does not fix generation errors.

How to reason it out
  1. Check what supervision exists: no labeled solutions.
  2. Check what feedback exists: an automatic pass rate.
  3. Map automatic scores to reinforcement fine-tuning with a reward function.
  4. Reject distillation because no teacher performs well on this task.

Exam tip: No labels but an automatic scorer: reinforcement fine-tuning with a reward function.

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02) — the lesson that teaches this.

Question 4ML Model and Foundation Model (FM) Development

A data science team is customizing a 70-billion-parameter open-weight model from SageMaker JumpStart for three separate tasks: claims triage, letter drafting and call summarization. GPU budget is tight, and the team wants to keep a separate variant per task without storing three full copies of the model. Which fine-tuning strategy should the team use?

Choose one.

  • a
    Parameter-efficient fine-tuning with LoRA, keeping one adapter for each task Correct

    LoRA freezes the base weights and trains small low-rank adapters, which cuts GPU memory and training cost; each task keeps a small adapter on one shared base model.

  • b
    Full-parameter fine-tuning per task, storing one complete model copy for each task

    Updating every weight needs the most GPU memory and compute, and it produces three full-size model copies, which the team wants to avoid.

  • c
    Domain adaptation fine-tuning per task on each task's unlabeled text

    Domain adaptation (continued pre-training) teaches vocabulary from unlabeled text; it does not teach task behaviour, and run per task it still produces full model copies.

  • d
    Full-parameter fine-tuning once on the combined data of the three tasks

    One combined model avoids three copies but still pays the full-parameter training cost and removes the separate per-task variants the team asked for.

The concept

Parameter-efficient fine-tuning (PEFT) such as LoRA trains a small set of added weights while the base model stays frozen.

Why that’s the answer

Two constraints decide it: limited GPU budget and a separate variant per task without three full copies. LoRA trains low-rank adapters, so training needs far less memory and compute than updating every weight, and each task's result is a small adapter that loads onto one shared base model. Full-parameter fine-tuning fails the budget and the storage constraint; a single combined model fails the per-task requirement; domain adaptation solves a different problem.

How to reason it out
  1. Note the budget constraint on training compute.
  2. Note the need for per-task variants without full copies.
  3. Recall that PEFT freezes the base and trains small adapters.
  4. Choose one adapter per task on a shared base model.

Exam tip: Many task variants on a tight GPU budget: LoRA adapters on one frozen base model.

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02) — the lesson that teaches this.

Question 5ML Model and Foundation Model (FM) Development

A telecom company wants to predict which subscribers will cancel next month. It has two years of history: 2 million labeled rows with 40 numeric and categorical columns such as plan type, monthly usage and support tickets. A churn probability for every subscriber is needed each night. Which modeling approach should the ML engineer choose?

Choose one.

  • a
    Prompt a foundation model in Amazon Bedrock with each subscriber's record

    A generative model reading one row per prompt is slow and costly for millions of rows and does not learn from the 2 million labeled outcomes.

  • b
    Train an Amazon Comprehend custom classifier on the subscriber records

    Comprehend custom classification is built for natural-language documents, not numeric and categorical feature columns.

  • c
    Run k-means clustering on the subscriber records to find churners

    Clustering groups records without using the labels, so it does not produce a churn probability for each subscriber.

  • d
    Train a gradient-boosted tree model such as XGBoost on Amazon SageMaker AI Correct

    Labeled tabular data with a binary outcome is the classic case for gradient-boosted trees, which are accurate, quick to train and cheap to score for millions of rows.

The concept

Structured, labeled, tabular prediction problems are best served by traditional supervised ML, not generative models.

Why that’s the answer

The data is tabular, labeled with the outcome, and large, and the output is a probability per subscriber every night. A supervised gradient-boosted tree model fits exactly this pattern. A foundation model is the tempting modern choice, but prompting it per row is expensive, slow and ignores the labeled history. Comprehend works on text, and clustering ignores the labels.

How to reason it out
  1. Identify the data type: labeled tabular rows.
  2. Identify the output: a probability per subscriber, in bulk.
  3. Map that to supervised tabular ML such as gradient-boosted trees.
  4. Reject generative and unsupervised options.

Exam tip: Labeled tabular prediction at volume: traditional supervised ML, not an FM.

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02) — the lesson that teaches this.

Question 6ML Model and Foundation Model (FM) Development

A lender is replacing its credit-approval scorecard. The regulator requires the lender to publish how much each input feature raises or lowers the score, as one fixed weight per feature that applies to every applicant. In validation, an XGBoost model scores 0.81 AUC and a logistic regression model scores 0.80 AUC. Which model should the ML engineer deploy?

Choose one.

  • a
    The 0.81 AUC XGBoost model, with SageMaker Clarify SHAP values reported for each decision

    SHAP values explain individual predictions after the fact and differ from applicant to applicant; they do not give one fixed published weight per feature.

  • b
    The 0.81 AUC XGBoost model, with its global feature-importance ranking published

    Feature importance says how much a feature is used, not the direction or size of its effect on the score, so it does not meet the disclosure requirement.

  • c
    The 0.81 AUC XGBoost model, retrained with monotonic constraints on each feature

    Monotonic constraints fix the direction of each feature's effect, but the size of the effect still varies with the other features, so there is no single published weight per feature.

  • d
    The 0.80 AUC model, trained with the SageMaker AI Linear Learner algorithm Correct

    The 0.80 AUC model is the logistic regression, and logistic regression is intrinsically interpretable: each feature has one coefficient that applies to every applicant, which is exactly the artefact the regulator asks for, at almost no accuracy cost.

The concept

Interpretability requirements can rule out more accurate models: intrinsically interpretable models such as linear or logistic regression expose their logic directly as coefficients.

Why that’s the answer

Two facts combine. The requirement is a fixed, published weight per feature, which is what a linear model's coefficients are, and the accuracy gap is small (0.81 vs 0.80 AUC). Each XGBoost option looks workable but misses part of the requirement: SHAP values from SageMaker Clarify are per-prediction attributions that vary by applicant, a global importance ranking has neither direction nor size, and monotonic constraints guarantee direction only, not a constant effect size. Only the logistic regression model gives one weight that applies to everyone.

How to reason it out
  1. Read the requirement literally: one fixed weight per feature for everyone.
  2. Recognise that as a linear model's coefficients.
  3. Weigh the accuracy difference: 0.01 AUC.
  4. Check each explainability add-on for XGBoost: per-prediction (SHAP), unsigned (importance) or direction-only (monotonic constraints).

Exam tip: A required global, fixed weight per feature means an intrinsically interpretable linear model; SHAP, importance and monotonic constraints do not provide it.

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02) — the lesson that teaches this.

What MLA-C02 domain 2 tests, topic by topic

The official exam guide breaks ML Model and Foundation Model (FM) Development into 3 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published MLA-C02 practice questions per topic in ML Model and Foundation Model (FM) Development
TopicWhat it coversQuestions
Choose appropriate modeling approaches for ML and AI solutionsExam guide task 2.1 (MLA-C02). Selecting foundation models in Amazon Bedrock by task requirements and performance; fine-tuning strategies for pre-trained FMs; comparing ML models, GenAI models, algorithms and solution templates (interpretability, domain-specific performance, latency); trade-offs between custom solutions, managed services, pre-trained models and FMs; choosing Retrieval Augmented Generation architecture patterns; trade-offs between model performance, training time, latency and cost; applying AWS AI services to business problems (Amazon Textract, Rekognition, Comprehend, Transcribe).38
Train, fine-tune, and customize models for ML and AI solutionsExam guide task 2.2 (MLA-C02). SageMaker AI built-in algorithms and common ML libraries; SageMaker AI script mode with supported frameworks; hyperparameter optimization (SageMaker AI automatic model tuning); reducing training time (early stopping, distributed training); preventing overfitting, underfitting and catastrophic forgetting; combining models (ensembles) for performance or cost; fundamental hyperparameters (epochs, steps, batch size); customization techniques for AI solutions (task-specific prompt engineering, fine-tuning); optimizing retrieval components and embedding models.38
Analyze and evaluate the performance of ML and AI systemsExam guide task 2.3 (MLA-C02). Reproducible experiments (MLflow on SageMaker AI, Amazon Bedrock evaluations, Bedrock Prompt Management); performance baselines and drift detection; comparing shadow variants with production variants; explaining model outputs; debugging convergence; evaluation techniques for traditional ML and GenAI models; human evaluation (human-in-the-loop, text generation quality assessment); NLP metrics (BLEU, ROUGE, BERTScore, semantic similarity); AI evaluation (output assessment, content quality, bias detection, LLM-as-a-judge); RAG monitoring and retrieval accuracy.38
Total114

Revise ML Model and Foundation Model (FM) Development before you drill it

Other MLA-C02 domains

ML Model and Foundation Model (FM) Development: your questions

ML Model and Foundation Model (FM) Development is domain 2 of the MLA-C02 exam guide and carries 24% of the scored content — the 2nd-heaviest of the 4 domains. On a 65-question paper that works out to roughly 16 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official MLA-C02 exam guide.