SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
Operating, Monitoring, and Securing ML and AI Solutions

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02)

18 min readMLA-C02 · Operating, Monitoring, and Securing ML and AI SolutionsUpdated

Monitoring ML and AI inference means continuously measuring whether models, foundation models and agents in production are still fast, error-free and accurate, and detecting when their input data or behaviour has drifted from what was approved. On AWS that splits into CloudWatch metrics and generative AI observability for Amazon Bedrock, SageMaker Model Monitor and Clarify for drift on SageMaker AI endpoints, production and shadow variants for live comparisons, Amazon Bedrock evaluations for FM and RAG quality, and agent traces plus AgentCore Evaluations for agent behaviour. This lesson explains which mechanism answers which question, what each one needs before it can work, and the near-twin settings that separate a working setup from one that silently misses the problem.

What you’ll learn
  • Use AWS/Bedrock metrics, model invocation logging and CloudWatch generative AI observability to monitor foundation model calls and agents.
  • Configure SageMaker Model Monitor data quality, model quality, bias drift and feature attribution drift monitors and read their violations.
  • Automate drift-triggered retraining and failure alerts with CloudWatch alarms, anomaly detection and Amazon EventBridge.
  • Run A/B tests with production variants and risk-free comparisons with shadow tests on SageMaker AI endpoints.
  • Choose the right Amazon Bedrock evaluation type and metrics for an FM or RAG use case, including scoring your own logged responses.
  • Detect agent tool failures, coordination failures and truncated streams, and automate agent quality scoring with AgentCore Evaluations.

Which monitoring tool answers which question?

Pick the tool from the signal that is failing. Task 4.1 covers four different kinds of problem, and each has its own AWS mechanism: operational health (is the model answering, how fast, with what errors), data and model drift (has the input or the model's behaviour moved since approval), live experiments (is version B better than version A on real users) and generative AI quality (are the answers, retrievals and agent decisions any good).

Question you need answeredMechanismNeeds
Latency, errors, throttles, tokens of Bedrock callsAWS/Bedrock CloudWatch metrics; Model Invocations view in CloudWatch generative AI observabilityNothing (metrics are automatic)
What exactly was the prompt and the responseBedrock model invocation loggingLogging turned on (off by default)
Have a SageMaker AI endpoint's inputs driftedModel Monitor data qualityData capture + baseline
Has accuracy, precision or recall droppedModel Monitor model qualityData capture + ground-truth labels
Has bias or feature reliance changedSageMaker Clarify bias drift / feature attribution driftData capture + Clarify baseline
Is version B better on live usersProduction variants (A/B)Two variants on one endpoint
Is version B safe, with no user exposureShadow variant / shadow testA copy of live traffic
Are FM or RAG answers goodAmazon Bedrock evaluationsA prompt dataset (optionally with your own responses)
Are agent sessions and tool calls goodAgent trace; AgentCore EvaluationsTrace enabled; AgentCore observability

Building dashboards, X-Ray and AgentCore Observability troubleshooting and token-cost tracking belong to Task 4.2; securing endpoints and Amazon Bedrock Guardrails belong to Task 4.3. This lesson covers detection and measurement.

Monitoring Amazon Bedrock with CloudWatch metrics and generative AI observability

Amazon Bedrock publishes runtime metrics to the AWS/Bedrock CloudWatch namespace for every model call, with no setup, and CloudWatch generative AI observability turns them into prebuilt views. The Model Invocations view shows invocation counts, token usage by model, latency percentiles, throttles and errors out of the box. Its per-request table (open one call, read its prompt and response) is filled only from model invocation logs delivered to CloudWatch Logs; logging only to Amazon S3 leaves that table empty. The Bedrock AgentCore view shows agent sessions, traces and error rates.

Choosing the right metric

MetricWhat it measuresTypical trap
InvocationLatencyTime until the whole response is completeMisses a slow start when total time is unchanged
TimeToFirstTokenDelay before the first token of a streaming response (ConverseStream, InvokeModelWithResponseStream)The right metric for "nothing appears for seconds"
InvocationThrottlesRequests throttled by the serviceRequests rejected for exceeding a quota are counted as client errors, not throttles
InvocationClientErrors / InvocationServerErrorsFailed requests by causeSee the ModelId dimension below
InputTokenCount / OutputTokenCountToken volumeLength, not topic or quality
EstimatedTPMQuotaUsageEstimated tokens-per-minute quota consumptionA capacity signal, not responsiveness
ModelInvocationLogsCloudWatchDeliveryFailure / ModelInvocationLogsS3DeliveryFailure / ModelInvocationLargeDataS3DeliveryFailureFailed attempts to deliver invocation logs to that destinationAlarm on the one matching your destination

The ModelId dimension undercounts early failures. Bedrock publishes runtime metrics both with and without a ModelId dimension. A request that fails before Bedrock resolves the target model (for example a malformed model or inference profile identifier) is counted only in the series without ModelId, so a per-model alarm can stay quiet during an outage. Alarm on the dimensionless metric too.

Agents hosted outside AgentCore Runtime

The AgentCore view reads OpenTelemetry spans and logs. To show an agent running on Amazon ECS, Amazon EKS, Lambda or elsewhere in the same place, enable CloudWatch Transaction Search in the account (the prerequisite for AgentCore observability) and instrument the agent with the AWS Distro for OpenTelemetry (ADOT) SDK, exporting spans and logs to CloudWatch.

Bedrock model invocation logging and Logs Insights analysis

Model invocation logging records the full request and response of every Bedrock call, and it is off by default. It is configured once per account and Region, with CloudWatch Logs, Amazon S3 or both as destinations; you choose which data types to include (text, image, embedding, video). CloudTrail is not a substitute: it records who called which API, never the prompt or response bodies.

  • Large bodies: a log event holds a JSON body up to 100 KB. Larger bodies are delivered only if you configure an S3 location for large data delivery; the log event then references the S3 object. The symptom of a missing large-data location is log events with metadata and token counts but no body for the biggest requests.
  • Attribution: each record carries identity.arn (the calling principal, which is identical for every application behind one shared role and session name) and requestMetadata, caller-supplied key-value tags you pass on Converse or InvokeModel calls. Adding an application tag to requestMetadata is a small code change that makes records filterable in Logs Insights with no new resources. Application inference profiles are the resource-based way to separate traffic and cost per application: each one is an AWS resource you create, and its ARN appears as the modelId.
  • Delivery health: if a role or log group breaks, calls still succeed but logs stop arriving. Alarm on the delivery-failure metric for each destination you use: ModelInvocationLogsCloudWatchDeliveryFailure, ModelInvocationLogsS3DeliveryFailure and, when a large-data location is configured, ModelInvocationLargeDataS3DeliveryFailure.

CloudWatch Logs Insights queries over the log group let you filter by model, metadata, latency or token count and read the exact prompts behind a quality complaint.

Detecting drift on SageMaker AI endpoints with Model Monitor and Clarify

SageMaker Model Monitor detects drift by comparing captured production traffic with a baseline on a schedule. Every monitor type needs data capture enabled on the endpoint (or batch transform job), which writes requests and responses to Amazon S3, plus a baseline and a monitoring schedule. Results include violation reports and, optionally, CloudWatch metrics for alarms. Model Monitor works only on SageMaker AI endpoints and batch transform; it cannot be attached to an Amazon Bedrock model.

MonitorDetectsBaseline fromNeeds labels?
Data qualityInput feature drift: distributions, missing values, types, schemaTraining data (statistics + constraints)No
Model qualityAccuracy, precision, recall, F1, RMSE and similar fallingModel predictions alongside true labels (e.g. validation set)Yes (ground truth)
Bias drift (Clarify)Post-training bias metrics for a facet (a sensitive attribute such as region or gender) movingClarify bias baselineDepends on metric
Feature attribution drift (Clarify)Which features drive predictions changing (SHAP ranking)Clarify SHAP baselineNo

Data quality violations

The data quality baseline job computes statistics and suggests constraints. Each run reports violations by check type:

  • data_type_check: a feature's inferred type differs from the baseline.
  • completeness_check: the share of non-null values fell below the baseline.
  • missing_column_check: a column in the baseline is absent.
  • baseline_drift_check: the distribution distance from the baseline exceeds the threshold. A shift in a feature's values that keeps its type, its non-null share and its column (for example an upstream change in how a value is scaled) trips only this check.

Bias drift versus feature attribution drift

Bias drift watches outcomes across groups: a post-training bias metric, such as the difference in positive prediction rates between two values of a facet, moving away from its baseline. Feature attribution drift watches how the model decides: Clarify computes SHAP attributions on live data and Model Monitor compares the feature-importance ranking with the baseline using an NDCG score. A model can keep accuracy and input distributions stable while either one moves. A Clarify pre-training bias job only describes the static training set and cannot detect production changes.

Model quality monitoring with delayed ground truth

Model quality monitoring measures real-world accuracy by merging captured predictions with ground-truth labels that you upload later, which is the only way to track precision or recall when outcomes arrive days or weeks after the prediction.

  1. Tag each request: pass a unique InferenceId on InvokeEndpoint. It is stored in the captured record, and your application keeps it with the business record (for example the transaction).
  2. Upload labels: when the outcome is known, write ground-truth records keyed by the same inference ID to the S3 location the model quality schedule reads.
  3. Merge and compare: the monitor joins captured predictions to labels by ID, computes metrics for the problem type and compares them with the baseline constraints, emitting CloudWatch metrics you can alarm on.

The baseline is different from a data quality baseline. Passing training features and labels to the data quality baselining job produces feature statistics only, never precision or F1 thresholds. A model quality baseline needs a dataset containing the model's predictions (and probabilities for threshold metrics) next to true labels, typically from running the current model on a held-out validation set, and a model quality baselining job told the problem type (BinaryClassification, MulticlassClassification or Regression) and which columns hold the inference and the ground truth. Unlabelled captured traffic cannot produce these metrics, and hand-written thresholds are not a measured baseline.

InferenceId only labels a request for this join; it does not route traffic.

Automating responses: drift pipelines and workflow anomaly detection

Detection becomes useful when it triggers action without polling code. Event-driven wiring through CloudWatch and Amazon EventBridge covers both drift-triggered retraining and failure alerts.

Drift-triggered retraining

Configure the monitoring schedule to publish its metrics to CloudWatch, create an alarm on the drift metric for the features that matter, and create an EventBridge rule matching the alarm's state change to ALARM with a SageMaker pipeline execution as its target. The alarm state is the trigger because it changes only when drift crosses the threshold; the job finishing says nothing about whether drift was found.

Failures in pipelines and jobs

SageMaker Pipelines emits execution and step status change events, and training, processing and transform jobs emit job state changes, to EventBridge. A rule on a Failed status with an Amazon SNS target notifies on-call within minutes. AWS Config evaluates resource configuration, not run outcomes.

Anomalies that are not failures

SituationUse
A metric with daily and weekly seasonality (invocations, latency)A CloudWatch alarm on an anomaly detection band; the model learns the pattern, no hand-tuned thresholds
A known error string in logsA metric filter + alarm
New, unpredictable log patterns while the job still succeedsCloudWatch Logs anomaly detection on the log group
Which clients contribute most trafficContributor Insights (ranking, not anomaly alerting)

A/B testing and shadow testing on SageMaker AI endpoints

Use production variants when users should receive the new model's answers in a controlled share; use a shadow variant when no user may see them. Both run on one endpoint and report CloudWatch metrics per variant.

Production variants (A/B)Shadow variant / shadow testDeployment guardrails (canary, linear)
Users get the new model's outputYes, for its weight shareNever; responses are captured for comparisonYes, during the shift
PurposeStanding comparison of business resultsValidate latency, errors and predictions on identical live trafficSafe, temporary rollout of an update (Domain 3)
Traffic controlVariant weightsTraffic sampling percentage of copied requestsShift steps with auto-rollback alarms

Controls you are expected to know

  • Weights: set initial weights per variant in the endpoint configuration (90/10 sends 10 percent of requests to the second variant).
  • UpdateEndpointWeightsAndCapacities changes weights (and instance counts) on the running endpoint in place with no new endpoint configuration. UpdateEndpoint also avoids downtime but requires a new endpoint configuration.
  • TargetVariant on InvokeEndpoint pins one request to a named variant, overriding the split for that request only (useful for QA). Setting it on every request replaces the endpoint's random split with client routing. CustomAttributes are passed to the container and do not route; InferenceComponentName targets inference components, not variants.
  • InvokedProductionVariant in the InvokeEndpoint response names the variant that served the request. CloudWatch has per-variant operational metrics only; any outcome measured by the application itself has to be joined to the variant that produced it, and this response field is what makes that join possible.

Shadow tests

A SageMaker AI shadow test copies a configurable percentage of live requests to the shadow variant (useful when the candidate runs on costly instances), returns only the production variant's responses, shows a console dashboard comparing both variants' latency and errors, and lets you promote the shadow variant to production at the end. Because it uses copies of real requests, a shadow test reflects the live input mix as it is today.

Detecting distribution change in generative AI traffic

For foundation models on Amazon Bedrock, detect input drift from the invocation logs, because Model Monitor data capture does not exist for Bedrock. Prompts are free text, so compare them in embedding space: a scheduled job embeds the recent logged prompts, compares their distribution (centroid distance, cluster proportions or a similar statistic) with launch-period embeddings, and publishes the result as a custom CloudWatch metric that can be alarmed on.

  • Token counts only measure prompt length; new topics can have the same length profile.
  • Re-running an evaluation on a fixed dataset of old prompts measures the model, not whether incoming questions changed.
  • Embedding drift detects that the traffic moved; evaluation (next section) tells you whether answer quality suffered as a result.

Amazon Bedrock evaluations: choosing the job type and metrics

Amazon Bedrock evaluations measure foundation model and RAG quality, and the job type follows from who judges and what is being judged.

Job typeJudgeMeasuresChoose when
Automatic (programmatic) model evaluationAlgorithmsAccuracy, robustness, toxicity for task types (text generation, summarization, Q&A, classification)Fast computed scores on your prompt dataset, no humans; it invokes the model
Model evaluation with LLM as a judgeA judge modelBuilt-in quality metrics (correctness, completeness, helpfulness, faithfulness, instruction following, coherence and responsible-AI metrics) plus custom metricsNuanced quality at scale without people
Human-based model evaluationYour own work team (private, Cognito-managed)Your rating methods (preference, Likert, thumbs)Subjective qualities only people can judge; compares up to two models side by side
RAG evaluation (retrieve-only or retrieve-and-generate)A judge modelRetrieval and grounded generationKnowledge bases or your own RAG system

Bring your own inference responses

Judge-based and RAG jobs accept a dataset that already contains each prompt's response, so Bedrock skips invocation and scores the answers as they were delivered. This is how you score logged production conversations, or a model hosted outside Bedrock (a SageMaker AI endpoint, a LangChain RAG stack), without generating new answers. Anything that invokes a model again scores fresh answers, not the ones users actually received.

Matching metrics to defects

Defect or needMetric
Do retrieved passages bear on the questionContext relevance (retrieve-only)
Do retrieved passages contain the reference factsContext coverage (retrieve-only; needs ground truth)
Answer states facts found in no retrieved passageFaithfulness
Citations point to passages that do not support the sentenceCitation precision
Supported sentences lack citationsCitation coverage
Answer is wrong or misses parts of the questionCorrectness / completeness (reference responses improve them)
An organisation-specific rubric (e.g. brand tone scored 1 to 5)A custom judge metric: a prompt with {{prompt}} and {{prediction}} variables and a rating scale

To compare chunking strategies, isolate retrieval with a retrieve-only job; retrieve-and-generate metrics mix in the generator's behaviour. Human jobs use a private work team managed through Amazon Cognito; Amazon Augmented AI (A2I) is a production review loop for individual predictions, not a model comparison.

Monitoring agents: tool failures, coordination failures and truncated streams

Agents fail in ways that request metrics do not see: a tool returns an error the agent politely explains, a supervisor routes to the wrong collaborator, or a stream ends early with HTTP 200. Each needs a specific signal.

Tracing what the agent did

Enable trace on InvokeAgent for an Amazon Bedrock agent. Trace events cover pre-processing, orchestration (the model's rationale, the action group invocation input with its parameters, and the observation it returned) and failure traces. With multi-agent collaboration, the supervisor's orchestration trace records each hand-off as structured fields: an invocation input of type AGENT_COLLABORATOR (collaborator name, alias ARN, input text) and an observation with the collaborator's output; trace parts also carry collaboratorName and callerChain. CloudTrail shows only that calls happened, Lambda logs miss the agent's side, and invocation logs hold the same story as raw prompt text.

Handled tool failures

Lambda's Errors metric counts only invocations that end in an unhandled error. A tool function that catches an exception and returns a message reports success, and the agent's own invocation succeeds too, so neither Lambda nor AWS/Bedrock/Agents error metrics move. To alert without changing what users are told, emit a custom metric for each handled failure (for example with CloudWatch embedded metric format log lines) and alarm on it.

Truncated streaming responses

With ConverseStream, the HTTP status is sent before tokens flow, so 200 never proves the answer finished. A complete stream ends with a messageStop event (carrying stopReason) followed by a metadata event with usage and latency.

SignalMeaningAction
stopReason: end_turnModel finished naturallyNone
stopReason: max_tokensOutput hit the maxTokens limitRaise maxTokens or continue the generation
stopReason: stop_sequenceA configured stop sequence matchedOnly possible if stopSequences are set
Other stop reasons (e.g. tool_use, guardrail_intervened, content_filtered, model_context_window_exceeded)The model paused for a tool call, or output was blocked or cut by a guardrail, filter or context limitHandle each explicitly; only end_turn (or an expected stop_sequence) means a natural finish
Exception event in the stream (modelStreamErrorException, throttlingException...)Failure after streaming startedHandle and retry
Stream closes with no messageStopAnswer cut offTreat as incomplete and retry

Automated agent quality: AgentCore Evaluations

Online evaluation continuously scores a sampled share of live sessions (a sampling rate you set; the default is 10 percent) from an AgentCore-observed agent's traces, using LLM-as-a-judge evaluators. On-demand evaluation scores the sessions you hand it when you call it. Built-in evaluators work at different levels, and the level must match the problem:

  • Session: Builtin.GoalSuccessRate (did the conversation achieve the user's goal).
  • Trace: Builtin.Helpfulness and similar per-reply qualities.
  • Tool: Builtin.ToolSelectionAccuracy and Builtin.ToolParameterAccuracy (right tool, right arguments) for silent delegation errors that raise no exception.

You can add custom evaluators. By default results are written as JSON events to /aws/bedrock-agentcore/evaluations/results/<config-id> (with session and trace IDs, so Logs Insights can list the lowest-scoring sessions) and scores are published as metrics in the Bedrock-AgentCore/Evaluations namespace, so a standard CloudWatch alarm pages on a falling score with no code.

Tip. Task 4.1 questions describe a model, foundation model, RAG system or agent in production and a symptom or requirement, then ask which monitor, metric, evaluation or configuration detects it. The deciding detail is usually a mid-scenario constraint: no ground-truth labels, labels arriving weeks later, no human reviewers, answers must be scored as delivered without re-invoking the model, no new AWS resources, users must never see the candidate's output, the agent's replies must not change, or as little custom code as possible. Expect near-twin options that differ in one component: data quality versus model quality versus bias drift versus feature attribution drift, production versus shadow variants, UpdateEndpoint versus UpdateEndpointWeightsAndCapacities, TargetVariant versus InvokedProductionVariant, InvocationLatency versus TimeToFirstToken, retrieve-only versus retrieve-and-generate metrics, online versus on-demand evaluation and session versus tool-level evaluators.

Key takeaways
  • Bedrock invocation logging is off by default; the Model Invocations view's per-request table needs it delivered to CloudWatch Logs, and bodies over 100 KB need an S3 large-data location.
  • TimeToFirstToken measures a slow start in streaming; InvocationLatency measures the whole response. Early failures appear only in Bedrock metrics without the ModelId dimension.
  • Every Model Monitor type needs data capture and a baseline; model quality also needs ground truth joined by InferenceId and a baseline built from predictions plus labels.
  • Data quality watches inputs, model quality watches accuracy, bias drift watches outcome gaps between groups, feature attribution drift watches the SHAP importance ranking.
  • Alarm on the drift metric, then let an EventBridge rule on the alarm state start the retraining pipeline; use anomaly detection bands for seasonal metrics and Logs anomaly detection for unpredictable log patterns.
  • Production variants expose users to the new model by weight; shadow variants never do. UpdateEndpointWeightsAndCapacities changes weights in place, TargetVariant pins a single request, InvokedProductionVariant attributes outcomes.
  • Bedrock judge-based and RAG evaluations can score your own logged responses without re-invoking the model; retrieve-only jobs isolate retrieval quality.
  • Handled tool errors are invisible to Lambda Errors; emit a custom metric. A stream without messageStop or with stopReason max_tokens is incomplete, whatever the HTTP status.
  • AgentCore online evaluation samples live sessions continuously; pick session, trace or tool-level evaluators to match the failure, and alarm on the Bedrock-AgentCore/Evaluations metrics.

Frequently asked questions

What is the difference between SageMaker Model Monitor data quality and model quality monitoring?

Data quality monitoring compares captured production inputs with statistics and constraints from the training data, reporting drift, missing values, type changes and missing columns without needing labels. Model quality monitoring merges captured predictions with ground-truth labels you upload later (joined by InferenceId) and tracks metrics such as accuracy, precision, recall or RMSE against a baseline built from predictions plus labels.

Can SageMaker Model Monitor watch a model on Amazon Bedrock?

No. Model Monitor works on data captured from SageMaker AI endpoints and batch transform jobs. For Bedrock, use model invocation logging as the record of prompts and responses, CloudWatch metrics for operational health, a scheduled job that compares prompt embeddings with a baseline for input drift, and Amazon Bedrock evaluations for answer quality.

When should I use a shadow test instead of an A/B test on SageMaker AI?

Use a shadow test when no customer may receive the new model's output: the shadow variant gets a copy of live requests (all or a sampled percentage), its responses are captured for comparison and never returned, and it can be promoted afterwards. Use production variants with weights for an A/B test when a share of users should receive the new model so business results can be compared.

How do I evaluate production answers from a model without calling it again?

Run an Amazon Bedrock judge-based model evaluation or RAG evaluation job with your own inference responses: the dataset includes each logged prompt and the response that was delivered, so Bedrock skips invocation and scores those exact answers. This also works for models and RAG systems hosted outside Bedrock.

Why does my Lambda Errors alarm stay at zero when agent tool calls are failing?

Lambda counts only invocations that end in an unhandled error. If the tool function catches the exception and returns an error message to the agent, the invocation succeeds, and the agent's own call succeeds too. Emit a custom metric for each handled failure, for example with CloudWatch embedded metric format, and alarm on that.

How can I tell that a Bedrock streaming response was truncated?

Check the end of the ConverseStream event stream. A stopReason of max_tokens in messageStop means the output limit was reached, an exception event inside the stream means a failure after streaming began, and a stream that closes with no messageStop at all is incomplete. The HTTP 200 status is sent before tokens flow and does not prove completion.

Which AgentCore evaluator detects an agent calling the wrong tool?

The tool-level built-in evaluators Builtin.ToolSelectionAccuracy (was the right tool chosen) and Builtin.ToolParameterAccuracy (were its arguments correct). Session-level Builtin.GoalSuccessRate and trace-level Builtin.Helpfulness score whole conversations or replies and do not isolate individual tool calls. Run them in an online evaluation to score sampled production sessions continuously.

Source

This lesson covers the "Operating, Monitoring, and Securing ML and AI Solutions" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.

Test yourself on this topic
Practice questions with full explanations.
Practice now

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?