SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
MLA-C02 · Domain 4

Operating, Monitoring, and Securing ML and AI Solutions practice questions

Operating, Monitoring, and Securing ML and AI Solutions is worth 24% of the MLA-C02 exam — the 2nd-heaviest of the 4 domains. Monitoring models, agents and data in production, optimizing infrastructure and inference cost, and securing ML and AI workloads and endpoints. 6 fully worked examples are further down this page, answers included.

Exam weight
24%
the 2nd-heaviest of the 4 domains
Questions
114
across 3 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Operating, Monitoring, and Securing ML and AI Solutions questions, fully explained

Questions from the MLA-C02 bank mapped to domain 4, with the answer key and the reasoning behind every option. None of them repeat the examples on the main MLA-C02 practice page.

Question 1Operating, Monitoring, and Securing ML and AI Solutions

A chat application streams answers from Amazon Bedrock by using the ConverseStream API. After a prompt template change, users complain that nothing appears on screen for several seconds after they press Enter, although the complete answer finishes in about the same total time as before. The ML engineer wants a CloudWatch alarm that tracks exactly this experience. Which metric should the alarm use?

Choose one.

  • a
    The InvocationLatency metric in the AWS/Bedrock namespace

    InvocationLatency runs until the last token arrives; users say total time has not changed, so this metric would barely move.

  • b
    The TimeToFirstToken metric in the AWS/Bedrock namespace Correct

    TimeToFirstToken measures the delay until the first token of a streaming response arrives, which is the wait users are describing.

  • c
    The OutputTokenCount metric in the AWS/Bedrock namespace

    Output token count reflects answer length, not how long the user waits before text starts to appear.

  • d
    The EstimatedTPMQuotaUsage metric in the AWS/Bedrock namespace

    This estimates tokens-per-minute quota consumption; it is a capacity signal, not a measure of perceived responsiveness.

The concept

For streaming Bedrock calls, TimeToFirstToken measures the wait before output starts; InvocationLatency measures until the last token.

Why that’s the answer

The users' complaint is the delay before the first text appears, while overall duration is unchanged. TimeToFirstToken is emitted for ConverseStream and InvokeModelWithResponseStream and measures exactly that gap. InvocationLatency covers the whole response and would hide the regression; token counts and quota estimates describe volume, not responsiveness.

How to reason it out
  1. Separate time-to-first-output from total response time.
  2. Note the call is streaming, so a first-token metric exists.
  3. Pick the metric that changed in the users' description.

Exam tip: Streaming responsiveness = TimeToFirstToken; total response time = InvocationLatency.

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.

Question 2Operating, Monitoring, and Securing ML and AI Solutions

An ML engineer built a CloudWatch alarm on InvocationClientErrors for each model, using the ModelId dimension. After a deployment that misconfigured the inference profile identifier in some requests, application logs show thousands of failed Bedrock calls, but the per-model alarm graphs show almost no client errors. What explains the gap?

Choose one.

  • a
    Requests rejected for exceeding a service quota are counted in InvocationThrottles, not under the ModelId dimension.

    It is the reverse: requests rejected for exceeding a quota are counted as client errors, not throttles. The failures here are bad identifiers anyway.

  • b
    Requests that fail input validation are published in the AWS/Bedrock/Agents namespace instead of by ModelId.

    The Agents namespace holds agent runtime metrics; failed bedrock-runtime calls are not moved there.

  • c
    Client error metrics per ModelId dimension come from model invocation logs, which are disabled by default.

    Runtime metrics are published whether or not invocation logging is enabled; logging adds request content, not metrics.

  • d
    Requests that fail before Bedrock resolves the target model publish metrics without the ModelId dimension. Correct

    Bedrock publishes runtime metrics both with and without ModelId; a request that fails before the model is known only appears in the dimensionless series.

The concept

AWS/Bedrock runtime metrics are published with and without the ModelId dimension; failures that happen before the model is determined only appear without it.

Why that’s the answer

The failures came from a bad model or inference profile identifier, so Bedrock could not attribute them to a model. Those requests are counted in the metric series without a ModelId dimension, so filtering by ModelId undercounts errors. To see all errors, alarm on the metric without a dimension (or both). Quota rejections count as client errors rather than throttles, and invocation logging has no effect on metric publication.

How to reason it out
  1. Note the cause: the request named a model that could not be resolved.
  2. Recall that Bedrock publishes runtime metrics with and without ModelId.
  3. Conclude that early failures only appear in the dimensionless series.

Exam tip: Alarm on Bedrock error metrics without the ModelId dimension if you need every failure counted.

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.

Question 3Operating, Monitoring, and Securing ML and AI Solutions

A retailer serves a demand-forecasting model from a SageMaker AI real-time endpoint. The ML engineer needs an hourly automated check of whether the live input features still look like the training data, including missing values, data types and value distributions, with results that can trigger CloudWatch alarms. The team has no ground-truth labels in production. What should the engineer configure?

Choose one.

  • a
    Enable data capture on the endpoint, create a baseline from the training data with SageMaker Clarify, and schedule a bias drift monitoring job.

    Bias drift tracks bias metrics across facets of the population, not missing values, types or general feature distributions.

  • b
    Enable data capture on the endpoint, create a baseline from the training data, and schedule a data quality monitoring job. Correct

    Data quality monitoring compares captured inputs with the baseline statistics and constraints (completeness, types, distribution drift) and can publish CloudWatch metrics.

  • c
    Enable data capture on the endpoint, create a baseline from the training data, and schedule a model quality monitoring job.

    Model quality monitoring compares predictions with ground-truth labels, which this team does not have, and it measures accuracy rather than input data.

  • d
    Export the captured requests each hour and generate a SageMaker Data Wrangler data quality and insights report for the team to review.

    Data Wrangler reports are an interactive data-preparation tool; this gives a manual review, not a scheduled check that drives alarms.

The concept

SageMaker Model Monitor data quality monitoring compares captured endpoint inputs with a baseline built from training data.

Why that’s the answer

The requirement is about inputs, not predictions, and there are no labels. A data quality monitor uses data capture plus a baseline (statistics and constraints computed from the training set) and runs on a schedule, reporting violations and CloudWatch metrics. Model quality needs labels, bias drift addresses fairness, and a Data Wrangler report is a manual analysis.

How to reason it out
  1. Identify what is monitored: input features, not accuracy.
  2. Note that no labels exist, which rules out model quality monitoring.
  3. Choose the Model Monitor type that checks inputs against a training baseline.

Exam tip: Inputs vs training data, no labels needed = Model Monitor data quality.

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.

Question 4Operating, Monitoring, and Securing ML and AI Solutions

A SageMaker AI endpoint predicts equipment failures from sensor readings, and a data quality monitoring schedule runs against a baseline built from the training data. A firmware update changes one sensor to report temperature in Fahrenheit instead of Celsius. The values are still valid numbers and no readings are missing. Which result should the ML engineer expect in the next monitoring report?

Choose one.

  • a
    A baseline_drift_check violation for the temperature feature Correct

    The unit change shifts the whole distribution of values; the distribution distance from the baseline exceeds the threshold and is reported as drift.

  • b
    A completeness_check violation for the temperature feature values

    Completeness compares the share of non-null values with the baseline; no readings are missing, so it should still pass.

  • c
    A data_type_check violation for the temperature feature

    The values are still numbers, so the inferred type matches the baseline.

  • d
    A missing_column_check violation for the temperature feature

    The column is still present in every record; nothing was removed from the schema.

The concept

Model Monitor data quality violations are typed: completeness, data type, missing/extra column, categorical values, and baseline drift (distribution distance).

Why that’s the answer

Nothing about the schema, types or nulls changed; only the values moved to a different scale. That is a distribution change, which the data quality monitor reports as baseline drift when the distance between the live and baseline distributions exceeds the configured threshold. The other checks look at structure and completeness, which still match.

How to reason it out
  1. List what changed: scale of values only.
  2. Rule out checks about nulls, types and columns.
  3. Map a pure distribution shift to the baseline drift check.

Exam tip: Same schema, shifted values = baseline drift, not a type or completeness violation.

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.

Question 5Operating, Monitoring, and Securing ML and AI Solutions

A payments company scores card transactions on a SageMaker AI endpoint with data capture enabled. Whether a transaction was really fraud is known only when a chargeback arrives, often weeks later. The ML engineer must track production precision and recall against the values measured at deployment and raise an alarm when they degrade. Which actions should the engineer take? (Select TWO.)

Choose TWO.

  • a
    Upload the chargeback labels, keyed by InferenceId, to the data quality monitor's baseline location so that it reports precision and recall.

    Data quality monitoring compares input feature statistics and constraints with its baseline; it never merges labels or computes classification metrics, whatever key the labels carry.

  • b
    Schedule a SageMaker Clarify bias drift monitor that compares the fraud rate for each merchant category with the training baseline.

    Bias drift tracks fairness metrics across facets; it does not report overall precision and recall degradation.

  • c
    Pass a unique InferenceId with each InvokeEndpoint request and store that ID with the transaction. Correct

    Model quality monitoring merges captured predictions with ground truth by inference ID, so every prediction needs an ID that the chargeback process can reference.

  • d
    Upload the chargeback labels, keyed by those inference IDs, to the S3 ground-truth location that a model quality monitoring schedule reads. Correct

    The model quality monitor merges these delayed labels with the captured predictions and computes precision and recall against its baseline.

  • e
    Store each chargeback in an online feature group and point the endpoint's data capture configuration at the feature group.

    Data capture writes to S3 and is not a label source; the feature group would not be merged with predictions by the monitor.

The concept

Model quality monitoring needs ground truth: predictions are captured with an inference ID, labels are uploaded later with the same ID, and the monitor merges them.

Why that’s the answer

Precision and recall need labels, so this is model quality monitoring, not data quality or bias drift. Because labels arrive late, the join key matters: send an InferenceId with each request (it is stored in the captured record) and upload ground-truth records with that same ID to the S3 prefix the model quality schedule reads. The monitor merges them, computes the metrics, compares them with the baseline and can emit CloudWatch metrics for alarms.

How to reason it out
  1. Precision/recall in production means labels are required.
  2. Pick the monitor type that uses labels: model quality.
  3. Work out how delayed labels are matched to predictions: an inference ID.
  4. Put the labels where the monitoring schedule reads ground truth.

Exam tip: Delayed labels + inference IDs + a model quality schedule give production accuracy metrics.

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.

Question 6Operating, Monitoring, and Securing ML and AI Solutions

A bank's credit-risk model runs on a SageMaker AI endpoint. Risk officers require an alert if the model begins basing its decisions on a different set of features than it did at approval, for example if a minor feature becomes one of the most influential. The input data distributions are monitored separately and currently look stable. Which monitoring should the ML engineer add?

Choose one.

  • a
    A SageMaker Model Monitor data quality schedule that uses a tighter distribution-distance threshold for each feature

    Tighter drift thresholds still watch input distributions, which are already monitored and stable; they say nothing about which features drive predictions.

  • b
    A SageMaker Clarify feature attribution drift monitoring schedule that uses the training data as the baseline Correct

    Feature attribution drift compares SHAP-based feature importance in production with the baseline and alerts when the importance ranking changes.

  • c
    A SageMaker Clarify bias drift monitoring schedule that uses the training data as the baseline

    Bias drift measures differences in outcomes between groups, not changes in which features the model relies on.

  • d
    A SageMaker Model Monitor model quality monitoring schedule with ground-truth labels from the loan outcomes

    Model quality measures accuracy-type metrics; a model can keep its accuracy while relying on different features.

The concept

Feature attribution drift (SageMaker Clarify with Model Monitor) tracks changes in feature importance between the baseline and live traffic.

Why that’s the answer

The officers care about how the model decides, not about input distributions or accuracy. Clarify's feature attribution drift monitor computes SHAP attributions on live data and compares the ranking with the baseline (Model Monitor uses an NDCG score to measure how far the ranking has moved). Bias drift is about outcomes across groups, and data or model quality would not reveal a change in feature reliance.

How to reason it out
  1. Ask what must be detected: a change in what drives predictions.
  2. Match 'feature influence' to SHAP feature attributions.
  3. Choose the Clarify monitor that compares attribution rankings over time.

Exam tip: Changing feature importance = feature attribution drift; changing outcomes between groups = bias drift.

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.

What MLA-C02 domain 4 tests, topic by topic

The official exam guide breaks Operating, Monitoring, and Securing ML and AI Solutions into 3 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published MLA-C02 practice questions per topic in Operating, Monitoring, and Securing ML and AI Solutions
TopicWhat it coversQuestions
Monitor ML and AI model inference and performanceExam guide task 4.1 (MLA-C02). Monitoring production model performance (Amazon CloudWatch generative AI observability, Amazon Bedrock Model Evaluation, drift detection pipelines); detecting anomalies or errors in data processing and inference; detecting data distribution changes; A/B testing in production; monitoring and automating agent performance and coordination (coordination failures, truncated streaming, tool failures); FM-specific performance monitoring (Amazon Bedrock evaluations).38
Optimize and manage ML and AI infrastructure costs and performanceExam guide task 4.2 (MLA-C02). Choosing inference instance families; troubleshooting tools (Amazon CloudWatch, Amazon Bedrock AgentCore Observability, AWS X-Ray); performance dashboards; capacity optimization for cost, performance and reliability; cost management tools and quotas; purchasing options; cost of FM inference in production; agent resource consumption; FM inference cost optimization; AI-specific cost patterns (token usage, embedding computation, vector database storage).38
Secure ML and AI workloads and model endpointsExam guide task 4.3 (MLA-C02). Securing CI/CD pipelines by scanning code and images (Amazon CodeGuru, Amazon Inspector); least-privilege access to ML and AI artifacts; IAM policies and roles for users and applications; monitoring, auditing, compliance and logging (AWS CloudTrail, AWS Config); troubleshooting security issues; VPCs, subnets and security groups to isolate ML and AI systems; identifying and mitigating ML and AI security risks; choosing credential types for FMs (Amazon Bedrock API keys, IAM credentials); safeguards and sensitive-data protection for responsible AI (Amazon Bedrock Guardrails).38
Total114

Revise Operating, Monitoring, and Securing ML and AI Solutions before you drill it

Other MLA-C02 domains

Operating, Monitoring, and Securing ML and AI Solutions: your questions

Operating, Monitoring, and Securing ML and AI Solutions is domain 4 of the MLA-C02 exam guide and carries 24% of the scored content — the 2nd-heaviest of the 4 domains. On a 65-question paper that works out to roughly 16 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official MLA-C02 exam guide.