SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
ML Model and Foundation Model (FM) Development

Evaluating ML and GenAI Models: Metrics, Drift, Explainability and RAG on AWS (MLA-C02)

24 min readMLA-C02 · ML Model and Foundation Model (FM) DevelopmentUpdated

Evaluating ML and AI systems means measuring how well a model, prompt or retrieval pipeline performs, explaining why it behaves as it does, and detecting when that performance changes in production. Task 2.3 of the MLA-C02 exam covers the whole loop on AWS: reproducible experiments with managed MLflow on SageMaker AI and Amazon Bedrock Prompt Management, performance baselines and drift detection, shadow testing, explainability with SHAP, debugging training that will not converge, choosing the right metric for traditional ML and for generated text, Amazon Bedrock evaluations (automatic, LLM-as-a-judge and human), bias detection, and RAG retrieval accuracy. Several SageMaker AI tools that used to anchor this topic (Model Monitor, Clarify and Debugger) are no longer open to new customers, so this lesson teaches the concepts first and shows both the existing-customer tool and the current replacement.

What you’ll learn
  • Make ML and prompt experiments reproducible with managed MLflow on SageMaker AI (runs, dataset inputs, code commits, tracking URI) and Amazon Bedrock Prompt Management versions and variants.
  • Build a performance baseline from training data and detect data drift and concept drift, including delayed ground truth, with Model Monitor (existing customers) or data capture, Evidently, MLflow, SNS and CloudWatch (new accounts).
  • Compare a new model with production using a SageMaker AI shadow test, knowing its endpoint prerequisites, data capture and promotion option.
  • Explain predictions with local SHAP values, global mean absolute SHAP and partial dependence plots, and measure post-training bias with DPPL, RD and DRR.
  • Diagnose divergence, vanishing and exploding gradients from loss curves and gradient histograms, and alert on training metrics without SageMaker Debugger.
  • Choose evaluation metrics for traditional ML (recall, precision, PR AUC, MAE, RMSE) and for text (BLEU, ROUGE, BERTScore, embedding similarity).
  • Select the right Amazon Bedrock evaluation (automatic, LLM-as-a-judge, human, retrieve-only or retrieve-and-generate RAG) and its metrics, rating methods and judge safeguards.

How do you make ML and prompt experiments reproducible on AWS?

An experiment is reproducible when you can rebuild its result from what you recorded: the code revision, the exact data, the parameters, the environment and the metrics it produced. On AWS that record is kept in managed MLflow on SageMaker AI for model training, and in Amazon Bedrock Prompt Management (plus repeatable Amazon Bedrock evaluations) for prompts and foundation models. Ad hoc records can hold numbers, but only an experiment tracker gives you run comparison, lineage and a model registry in one place.

Managed MLflow on SageMaker AI

SageMaker AI hosts open-source-compatible MLflow for you (a serverless MLflow App or a tracking server), so there is nothing to patch or scale. Each run belongs to an experiment and records parameters (mlflow.log_param), metrics (mlflow.log_metric), artifacts such as plots and the model itself, and tags. mlflow.autolog() captures hyperparameters, metrics and the trained model for common libraries automatically. The MLflow UI then compares runs side by side, and the MLflow Model Registry holds the chosen model as a registered, versioned model (how registered versions are approved and released is Task 3.3).

  • Connecting from a training job. A notebook in SageMaker AI Studio may already be wired to MLflow, but a training job is a fresh container. The script must call mlflow.set_tracking_uri() with the tracking server's (or App's) ARN, and the job needs the sagemaker-mlflow plugin (for example in requirements.txt) so MLflow can authenticate with the execution role's IAM permissions. If either is missing, MLflow silently falls back to local file storage inside the container, which is discarded with the container, so the job still succeeds while the runs are lost. Keep the channels apart: training-job metric definitions feed CloudWatch, and MLflow receives only what the script logs to the tracking server.
  • Logging what you need to rebuild a model. Parameters and metrics describe a run but cannot recreate it. Log the training data as an MLflow dataset input (mlflow.log_input), which records its source location (for example an S3 URI) and a content digest so you can tell whether the data changed, and tag each run with the Git commit ID of the code it executed. Data version + code version + parameters + environment is what makes a run reproducible.

Amazon Bedrock Prompt Management

Prompt Management stores a prompt (template text, variables, model and inference settings) as a managed resource with two kinds of state:

FeatureWhat it isUse it to
Working draftThe editable copy you save as you iterateExperiment; never point production at it, because every save changes what runs
VersionAn immutable, numbered snapshot of the draft with its own ARNShip a fixed prompt; roll back by calling an earlier version's ARN
VariantAn alternative configuration of the same prompt (different wording, model or settings) in the prompt builderRun the alternatives on the same input and compare outputs side by side before choosing

A version cannot be edited. To fix a released prompt, correct the draft, create a new version from it and repoint the application at the new version's ARN; the old version stays intact as the rollback target. Deleting a version breaks any caller that uses its ARN. For scored, repeatable comparisons of prompts or models across a whole dataset, run an Amazon Bedrock evaluation job (covered below) on a fixed prompt dataset, so the only thing that changes between runs is the thing you are testing.

How do you build a performance baseline and detect drift?

A performance baseline is the reference a production model is compared with: statistics and constraints computed from the training data, and quality metrics (accuracy, recall, RMSE and so on) measured on held-out data before release. Drift detection then compares live traffic and live outcomes with that baseline on a schedule and alerts when they move apart. Build the baseline from the data the model learned from, not from recent production traffic, which may already have drifted.

Kind of driftWhat changesHow you detect it
Data (feature) driftThe distribution of inputs moves away from the training dataCompare captured requests with training-data statistics, per feature (for example population stability index or a Kolmogorov-Smirnov test)
Concept driftThe relationship between inputs and the correct label changes; inputs can look exactly the sameJoin ground-truth labels to logged predictions and track quality metrics against the baseline
Feature attribution driftThe features the model relies on, measured by attributions such as SHAP, shiftCompare live attributions with training-time attributions
Bias driftFairness metrics across groups move over timeCompute the same bias metrics per segment on live predictions and alert on gaps

Delayed ground truth is the hard case. If fraud labels arrive weeks later, a new fraud pattern that looks like normal spending leaves every input-distribution metric flat while recall falls. Only model quality monitoring catches it: log each prediction with an inference ID, join the late labels by that ID, compute recall (or another quality metric) and alarm when it drops below the baseline. The principle: a monitor can only detect a change in what it measures, so a change confined to the input-to-label relationship needs a monitor on outcomes. Retraining is a response to detected drift, never a substitute for detecting it.

Which tools to use

Amazon SageMaker Model Monitor is no longer open to new customers (existing customers keep using it; no new features are planned). For an account that already has it, the workflow is: enable data capture on the endpoint (requests and responses written to Amazon S3), run a baselining job on the training dataset to produce statistics.json and suggested constraints.json, then create a monitoring schedule. Its monitor types are data quality, model quality (which merges ground-truth labels with captured predictions), bias drift and feature attribution drift; violations surface in reports and CloudWatch metrics.

For a new account, AWS's documented replacement is built from open building blocks: endpoint data capture, a scheduled job (EventBridge with Lambda, or a SageMaker AI Pipeline step) that runs Evidently drift and quality presets against the training baseline, results logged to a SageMaker AI MLflow App next to the training metrics, Amazon SNS notifications when a drift or quality check breaches its threshold, optional QuickSight dashboards, and Amazon CloudWatch for endpoint metrics (latency, errors, utilization) and anomaly-detection alarms on custom metrics. Keep the two layers distinct: infrastructure metrics say whether the endpoint is healthy, while data and quality checks say whether the model is still right.

How do shadow variants differ from production variants?

A shadow variant receives a copy of live requests sent to a SageMaker AI endpoint, but its responses are never returned to callers; a production variant serves real responses to the share of traffic its weight assigns. Use a shadow test when a new model must be judged on real traffic with zero customer impact; use weighted production variants (A/B testing, covered with production monitoring in Task 4.1) when you are ready for some customers to receive the new model's answers.

Shadow variant (shadow test)Production variant with a traffic weightBlue/green canary update
Who sees the new model's outputNobody; only the production response goes backThe weighted share of customersThe canary share, then everyone
PurposeCompare latency, errors and predictions on live traffic before switchingMeasure business impact of real answersSafely roll out a model already chosen
TrafficA sampled copy of production requestsSplit between variants by weightShifted from old fleet to new fleet

Running a shadow test in SageMaker AI:

  • Endpoint type matters. Shadow tests need an instance-based real-time endpoint. They are not supported on serverless inference, asynchronous inference, multi-model or multi-container endpoints, AWS Marketplace containers or Inf1 instances. A model hosted on an unsupported endpoint type has to be on an instance-based real-time endpoint before a shadow test can run.
  • Compare per request with data capture. Aggregate metrics describe each variant's latency and errors; comparing what the two models actually predicted needs per-request records. Enabling data capture for the test writes requests and both variants' responses to S3 for a request-by-request comparison. SageMaker AI copies the traffic to the shadow variant itself, so callers change nothing.
  • Promote the winner. When you complete the test you can choose to deploy the shadow variant, which replaces the production variant on the same endpoint, so clients keep calling the same endpoint name with no changes.

What a shadow test adds over any offline evaluation is measurement under real, concurrent production load: live latency, error rates and the actual request mix.

How do you explain model outputs and measure bias in predictions?

Explaining a model means attributing its outputs to its inputs, either for one prediction (local explanation) or across a whole dataset (global explanation). The standard tool is SHAP (Shapley additive explanations): for each prediction it assigns every feature a signed contribution relative to a baseline, and the contributions add up to the difference between the prediction and the baseline prediction.

Question being askedTechnique
Why did this applicant get this decision; which of their values mattered most?SHAP values for that one prediction (local attributions)
Which features drive the model most, ranked, across the whole portfolio?Mean absolute SHAP value per feature across the dataset (global importance)
How does the average prediction change as one feature moves across its range?Partial dependence plot (PDP) for that feature
How well does the model classify overall?Not an explanation question: confusion matrix and metrics

Taking the mean of absolute SHAP values matters: positive and negative contributions would otherwise cancel. Only local attributions answer a per-person question; any global ranking, however it is computed, describes the model as a whole. A PDP sets one feature to each value in turn for every row and averages the predictions, so it answers "what happens if we change X" questions such as a price change. Explanation techniques describe the model's behaviour; statistics computed on the data alone describe the data, not what the model learned.

SageMaker Clarify is no longer open to new customers (existing customers keep it). New accounts compute the same things with the open-source shap library in a notebook or processing job, bias formulas in pandas or scikit-learn, and fmeval or Amazon Bedrock evaluations for foundation models. The concepts and metric names are the same either way.

Post-training bias metrics for predictive models

Post-training bias metrics compare the trained model's predictions across groups (facets). They differ in what they condition on, which is why equal overall approval rates can hide unequal treatment:

  • DPPL (difference in positive proportions in predicted labels): the gap in predicted approval rates between groups, regardless of true outcome.
  • RD (recall difference): the gap in true positive rate, that is, among people whose true outcome was positive (for example, who repaid), how often each group was approved.
  • DRR (difference in rejection rates): among predicted rejections, how often the rejection was correct (true negatives) in each group.
  • DPL (difference in proportions of labels) is a pre-training metric on the training labels, not on model predictions; pre-training bias is part of data preparation (Task 1.3).

How do you debug model convergence problems?

A model fails to converge when its training loss does not settle at a low value: it diverges to NaN, oscillates, or stays flat. Read the loss curve and the gradients first, then change the one thing the symptom points to. (Overfitting, where training loss is fine but validation loss rises, is a generalization problem covered with training in Task 2.2.)

SymptomLikely causeFirst fix
Loss drops, then swings widely and becomes NaN (data already checked)Learning rate too highLower the learning rate, optionally with warmup or a decay schedule
Loss nearly flat from the start; gradients near zero in early layers but normal in late layersVanishing gradients from saturating activations (sigmoid, tanh) in a deep networkUse ReLU-family activations with He initialization; batch normalization and residual connections also help
Gradient norms spike; weights become huge or NaNExploding gradients (common in deep or recurrent networks)Clip gradients to a maximum norm; lower the learning rate
Loss plateaus at a high value for training and validationLearning rate too low, too little capacity or poor featuresRaise the learning rate or use a schedule; increase capacity

The principle: match the fix to the mechanism. Regularization targets overfitting, clipping caps gradients that are too large, and activation and initialization choices address gradients that shrink toward zero. A change aimed at a different mechanism leaves the symptom in place.

Tools for watching and inspecting training

SageMaker Debugger is no longer open to new customers (existing customers keep its built-in rules, such as loss-not-decreasing and vanishing-gradient). For a new account:

  • Alert: publish validation loss as a training job metric with a metric definition (a regular expression that SageMaker AI applies to the job's logs), then create an Amazon CloudWatch alarm on it.
  • Inspect: write TensorBoard summaries (scalars plus per-layer weight and gradient histograms) to the job's output in S3 and open them in TensorBoard on SageMaker AI.
  • Compare runs: log the loss curves to MLflow so a diverging run can be compared with a healthy one and its parameters.

Choose signals that measure learning (loss, gradient and weight statistics). Resource metrics and endpoint monitoring answer other questions: whether hardware is busy, or whether a deployed model is drifting.

Which metrics should you use to evaluate traditional ML models?

Choose the metric that matches the cost of each kind of error in the business problem; a single headline number such as accuracy can rank models wrongly. For classification, start from the confusion matrix (true and false positives and negatives); for regression, decide how much large errors should count.

MetricWhat it measuresUse when
AccuracyShare of all predictions that are correctClasses are balanced and errors cost the same; misleading with rare positives (a model that always predicts the majority class scores 98% when only 2% of cases are positive)
Recall (sensitivity, true positive rate)Share of actual positives caughtMissing a positive is expensive (fraud, disease) and reviewers can absorb false alarms
PrecisionShare of flagged items that are truly positiveFalse alarms are expensive or capacity to review is limited
F1 scoreHarmonic mean of precision and recall at one thresholdYou need one number that balances both at a chosen threshold
ROC AUCRanking quality across all thresholds, using true and false positive ratesClasses are reasonably balanced; under heavy imbalance it stays high for most models and stops separating them
PR AUC (average precision)Area under the precision-recall curve across all thresholdsPositives are rare and the threshold is not chosen yet
SpecificityShare of actual negatives correctly rejectedThe cost sits in the majority class
MAEAverage absolute error; every unit counts the sameCost grows linearly with error; robust to a few outliers
RMSESquare root of mean squared error; large errors weigh moreA few large misses are disproportionately costly

Two habits decide most metric choices. First, under class imbalance, report minority-class recall and precision (or PR AUC) instead of accuracy or specificity. Second, a threshold metric such as F1 at 0.5 is only meaningful once the operating threshold is known; until then compare models with a threshold-free metric. For regression, a model can have a lower RMSE but a higher MAE than another, when a few huge errors are pulling RMSE up. If each unit of error costs the same and the spikes are handled elsewhere, the lower MAE wins. The gap between RMSE and MAE describes how errors are distributed, not whether a model over- or underfits.

How do BLEU, ROUGE, BERTScore and semantic similarity differ?

These are reference-based NLP metrics: each scores generated text against one or more human-written reference texts, but they measure different things. BLEU and ROUGE count shared word sequences (n-grams); BERTScore and embedding similarity compare meaning through embeddings, so they credit correct paraphrases.

MetricHow it scoresTypical useWeakness
BLEUPrecision-oriented: share of the output's n-grams (up to 4-grams) found in the reference, with a brevity penalty for short outputMachine translation, the field's standard metricPenalizes valid synonyms and rephrasing
ROUGE (ROUGE-N, ROUGE-L)Recall-oriented: share of the reference's n-grams (ROUGE-N) or longest common subsequence (ROUGE-L) found in the outputSummarizationSame wording dependence: a faithful paraphrased summary scores low
BERTScoreMatches tokens of output and reference by cosine similarity of contextual embeddings from a BERT-style modelSummaries and answers where good outputs reword the referenceNeeds an embedding model; still needs references
Semantic (embedding) similarityCosine similarity between whole-text embeddings of output and referenceCheap, deterministic paraphrase-tolerant scoring of short answersMeasures closeness of meaning, not factual detail or style
PerplexityHow surprised a language model is by a textFluency or language-model qualityNot a comparison against a reference

Choose by what "good" looks like. If correct outputs share wording with the reference, as in translation, BLEU is standard; for summaries, ROUGE. If correct outputs are routinely paraphrased ("cancel your subscription" versus "end your plan"), n-gram scores fall even when editors approve, and you move to BERTScore or embedding similarity. Embedding similarity with a fixed embedding model is also repeatable and inexpensive, unlike a judge model, whose scores depend on another model's judgment. Classification F1 is a different metric from the token-overlap F1 used for short question answering.

Which Amazon Bedrock evaluation job should you run?

Pick the Amazon Bedrock evaluation job by who or what does the scoring: a fixed algorithm (automatic), a judge model (LLM-as-a-judge, and RAG evaluations), or people (human-based). Each job scores model or RAG responses on a prompt dataset you supply, or on built-in datasets.

Job typeWho scoresMetricsChoose it when
Automatic (programmatic)Fixed algorithmsAccuracy, robustness and toxicity, by task type: general text generation, summarization (accuracy by BERTScore), question and answer (accuracy by F1 against reference answers), text classificationScores must be repeatable and algorithmic, with no model judgment or people, for example regression tests after each prompt change
LLM-as-a-judgeAn evaluator (judge) model you chooseBuilt-in quality and responsible-AI metrics, plus custom metrics; each score comes with the judge's written explanationSubjective qualities (helpfulness, completeness, following instructions) at scale, with no human raters
Human-basedYour own work teamYour own metrics and rating methodsOnly experts can judge quality, or human preference is the deciding signal
RAG evaluationA judge modelRetrieval and generation metrics (see the RAG section)Evaluating a knowledge base or your own RAG pipeline

Automatic jobs can use built-in datasets, such as RealToxicityPrompts for toxicity, or your own prompt dataset with reference answers. A guardrail is not an evaluation: it filters content at runtime and produces no per-prompt scores.

Evaluating models that are not on Bedrock. Bedrock evaluations can score responses you generated elsewhere. Run the prompts through your model (for example, a fine-tuned LLM on a SageMaker AI endpoint), put the outputs in the prompt dataset as your own inference responses, and the job scores them alongside or against a Bedrock model, without moving or re-hosting your model. The job itself generates responses only from models it can invoke in Bedrock, so bringing your own responses is how any other model is included.

Human evaluation: human-in-the-loop quality assessment

A human evaluation job sends each prompt and the responses of up to two models (or inference sources) to a work team you manage, so confidential prompts are seen only by your own people, such as attorneys or medical reviewers. You define the metrics, write instructions, and set the number of workers per prompt so that several reviewers rate each item and one reviewer's bias is diluted; all ratings land in one report. The rating method decides what kind of judgment you collect:

Rating methodWhat the worker does
Choice buttonsPicks which of two side-by-side responses is preferred, nothing more
Likert scale, comparisonCompares two responses and rates how strongly one is preferred
Likert scale, individualRates one response on its own scale
Thumbs up or downJudges one response as acceptable or not
Ordinal rankRanks responses in order

Use human review where judgment is genuinely expert or subjective, and use it to calibrate automated judges. A human evaluation job is rated only by its human work team; when you want both kinds of score, run a separate LLM-as-a-judge job.

How do you assess GenAI output with LLM-as-a-judge and detect bias?

LLM-as-a-judge evaluation uses a separate evaluator model to score each response against named criteria and explain the score, so subjective quality can be measured at scale. Select metrics by the failure you need to track, because each built-in metric targets one failure mode.

Built-in metricCatches
FaithfulnessClaims not supported by the context given in the prompt (hallucination against the source)
CompletenessAnswers that address only part of a multi-part question
CorrectnessFactually wrong answers (can use a reference answer)
Helpfulness, Relevance, CoherenceUnhelpful, off-topic or illogical answers
Following instructionsIgnoring explicit instructions in the prompt
Professional style and toneTone and style problems, not organization-specific rules
HarmfulnessHarmful content in general
StereotypingGeneralized statements about people based on group membership such as gender, age or nationality
RefusalDeclining to answer

Custom metrics cover rules no built-in metric knows, such as approved product names or a required disclosure. You write the judge prompt that states the rules and a defined rating scale, and the job applies it to every response. Built-in metrics apply generic criteria; organization-specific rules reach the judge only through a custom metric.

Making a judge trustworthy

Judge models have systematic biases: self-preference (favouring outputs from their own model family), position bias (favouring the first of two answers) and verbosity bias (favouring longer answers). If the judge consistently disagrees with a blind expert review, the evaluation is not trustworthy. Use a judge from a different model family from the candidates, and calibrate it by scoring a sample with both the judge and human experts and checking agreement. A systematic bias is fixed by independence and calibration; a change that keeps the same judge family or only adds randomness leaves the bias in place.

Bias detection and content-quality validation

For generative output, bias is measured on the text: run an LLM-as-a-judge job with the Stereotyping metric on representative prompts to get a scored report per prompt, and use Harmfulness for harmful content more broadly. Evaluate on prompts representative of your own application: a generic benchmark describes the model in general, not your assistant's behaviour. Guardrails are the runtime control that blocks content; they complement evaluation but produce no evaluation report. For predictive models, bias is measured with the post-training metrics (DPPL, RD, DRR) covered in the explainability section.

How do you evaluate and monitor RAG retrieval accuracy?

RAG quality is evaluated in two layers, and the metrics tell you which layer is failing: retrieval (did the knowledge base return the right passages?) and generation (did the model use them faithfully?). Amazon Bedrock RAG evaluations provide a job type for each, for Amazon Bedrock Knowledge Bases or for your own RAG pipeline (by bringing your own retrieved passages and responses).

Evaluation typeMetricsAnswers
Retrieve onlyContext relevance (are retrieved passages relevant to the question?); context coverage (do they contain the information in the ground-truth answer? needs reference answers)Is the knowledge base the problem?
Retrieve and generateCorrectness, completeness, helpfulness, logical coherence, faithfulness, citation precision, citation coverage, plus harmfulness, stereotyping and refusalIs the final answer good, grounded and properly cited?

Localizing a fault. Before touching prompts, run a retrieve-only job: low context relevance or coverage points at chunking, embeddings, search type or the number of results (retrieval tuning is Task 2.2; knowledge base configuration and reranking are Domain 3). If context relevance and coverage are high but faithfulness is low, retrieval is fine and the generator is adding unsupported claims; fix generation, for example by instructing the model to answer only from the retrieved passages and to cite them, or by changing the model. The principle: fix the failing layer and leave the working one alone.

Citation metrics catch two different citation failures:

  • Citation precision: of the citations given, how many actually support their statement. Low when citations point at the wrong passages.
  • Citation coverage: how much of the response is supported by its citations. Low when claims are grounded in the retrieved text but carry no citation.
  • Faithfulness asks whether claims are grounded in the retrieved context at all, cited or not.

In production, keep the same evaluation dataset and rerun these jobs after each change to the knowledge base, embedding model or prompt, and track the scores over time with the other experiment results so that a drop in retrieval accuracy is caught before users report wrong answers. Runtime observability of live GenAI traffic in CloudWatch belongs to monitoring (Task 4.1).

Tip. Task 2.3 questions are scenarios that describe a model, prompt or RAG system that has to be compared, explained, debugged or monitored, usually with one constraint in the middle of the story: customers must see only production responses, the account is new (so tools closed to new customers are out), labels arrive weeks late, no human raters are available, scores must be repeatable, data is confidential, or the model must stay where it is hosted. Expect near-twin options that differ in one detail: a baseline from training data versus recent traffic, a retrieve-only versus retrieve-and-generate job, citation precision versus citation coverage, choice buttons versus Likert comparison, Stereotyping versus Harmfulness, local SHAP versus mean absolute SHAP versus a partial dependence plot, BLEU versus ROUGE versus BERTScore, MAE versus RMSE. Work out which component or failure mode the symptoms point to before choosing the metric or tool, and check whether the scenario needs an evaluation report or a runtime control.

Key takeaways
  • Reproducible runs need the data version (MLflow dataset input with a digest) and the code version (Git commit tag), not only parameters and metrics; a training job reaches managed MLflow only with the tracking server ARN as tracking URI and the sagemaker-mlflow plugin.
  • In Bedrock Prompt Management, iterate on the draft, compare alternatives as variants, and ship immutable numbered versions by ARN; fix a prompt by creating a new version, never by editing or deleting the old one.
  • Baselines come from the training data. Input-distribution checks catch data drift; only ground-truth quality metrics, joined by inference ID when labels arrive late, catch concept drift.
  • SageMaker Model Monitor, Clarify and Debugger are closed to new customers: new accounts use data capture with Evidently, MLflow and SNS/CloudWatch for drift, the SHAP library and pandas or scikit-learn bias formulas for predictive models, fmeval or Bedrock evaluations for foundation models, and metric definitions, CloudWatch alarms and TensorBoard for training.
  • A shadow variant sees live traffic but never answers customers; it needs an instance-based real-time endpoint, data capture compares responses, and completing the test can deploy the shadow variant.
  • Local explanation means one prediction's SHAP values; global ranking means mean absolute SHAP; how the prediction changes as one feature moves is a partial dependence plot.
  • Divergence to NaN: lower the learning rate. Vanishing gradients: ReLU with He initialization. Exploding gradients: clip.
  • With rare positives, ignore accuracy and compare recall, precision or PR AUC; linear error cost means MAE, costly large misses mean RMSE.
  • BLEU for translation, ROUGE for summaries, BERTScore or embedding similarity when correct answers are paraphrased.
  • Bedrock evaluations: automatic for repeatable algorithmic scores, LLM-as-a-judge (with custom metrics and an independent, human-calibrated judge) for subjective quality, human jobs with your own team for expert or confidential review, retrieve-only RAG jobs to test retrieval and retrieve-and-generate jobs for faithfulness and citations.

Frequently asked questions

What replaces SageMaker Model Monitor for a new AWS account?

SageMaker Model Monitor is no longer open to new customers, so AWS documents a replacement built from open components: enable data capture on the SageMaker AI endpoint, run a scheduled job (for example EventBridge with Lambda) that computes drift and quality with Evidently against the training-data baseline, log the results to a SageMaker AI MLflow App, alert through Amazon SNS, and use Amazon CloudWatch for endpoint metrics and anomaly-detection alarms. Existing Model Monitor customers can keep using it.

What is the difference between data drift and concept drift?

Data drift is a change in the distribution of a model's inputs compared with its training data, and it is detected by comparing captured requests with training-data statistics. Concept drift is a change in the relationship between inputs and the correct output, so inputs can look unchanged while predictions get worse. Concept drift is detected only by joining ground-truth labels to logged predictions and tracking quality metrics such as recall against the baseline.

When should I use a shadow variant instead of a production variant on SageMaker AI?

Use a shadow variant when a new model must be tested on live traffic without any customer receiving its responses: it gets a copy of requests and its responses are discarded or captured for comparison. Use a weighted production variant when you are ready for a share of customers to receive the new model's answers, as in an A/B test. Shadow tests require an instance-based real-time endpoint, not serverless, asynchronous, multi-model or multi-container endpoints.

What is the difference between BLEU, ROUGE and BERTScore?

BLEU and ROUGE both count overlapping n-grams between generated text and a reference: BLEU is precision-oriented and standard for machine translation, while ROUGE is recall-oriented and standard for summarization. BERTScore compares texts using contextual embeddings, so it credits correct paraphrases and synonyms that n-gram metrics score as wrong. Use BERTScore or embedding similarity when good outputs routinely reword the reference.

What evaluation types does Amazon Bedrock offer?

Amazon Bedrock offers automatic (programmatic) model evaluations that compute algorithmic metrics such as accuracy, robustness and toxicity; LLM-as-a-judge evaluations in which an evaluator model scores responses on built-in or custom metrics and explains each score; human-based evaluations rated by your own work team; and RAG evaluations, either retrieve-only or retrieve-and-generate. You can also bring your own inference responses to evaluate models hosted outside Bedrock.

How do you make an LLM-as-a-judge evaluation reliable?

An LLM judge can show self-preference toward its own model family, position bias and verbosity bias. Make it reliable by choosing a judge from a different model family from the models being compared, giving it clear criteria and a defined rating scale (a custom metric when the criteria are your own rules), and calibrating it by scoring a sample with both the judge and human experts and checking that they agree.

Which metrics show whether a RAG knowledge base is retrieving the right content?

A retrieve-only RAG evaluation in Amazon Bedrock measures retrieval directly with context relevance, which checks whether retrieved passages are relevant to the question, and context coverage, which checks whether they contain the information in the reference answer and therefore needs ground truth. Generation problems are measured separately in a retrieve-and-generate evaluation with faithfulness, correctness, completeness, citation precision and citation coverage.

How do you explain an individual prediction from a tabular model on AWS?

Compute SHAP values for that prediction: each feature gets a signed contribution that shows how much it pushed the prediction up or down relative to a baseline, which is what a per-customer explanation such as a loan decline notice needs. New AWS accounts compute SHAP with the open-source shap library because SageMaker Clarify is no longer open to new customers. For a global ranking, average the absolute SHAP values across the dataset.

Source

This lesson covers the "ML Model and Foundation Model (FM) Development" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.

Test yourself on this topic
Practice questions with full explanations.
Practice now

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?