SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
MLA-C02 last-day review

MLA-C02 study guide — every exam topic on one page

120 key facts across 4 exam domains — every topic on the MLA-C02 blueprint, with the exam pattern behind each. A condensed cheat sheet, distilled from the full MLA-C02 revision notes. Skim it the week of your exam.

Updated

Data Preparation for ML and AI

28% of the exam

Collecting and Storing Data for ML and AI on AWS (MLA-C02)

  • File mode downloads the whole channel before training; FastFile streams with file semantics and no script change; Pipe needs pipe-reading code. ShardedByS3Key splits data across instances, FullyReplicated copies all of it to each.
  • EFS and FSx for Lustre channels need VPC configuration; FSx for Lustre is the throughput choice for repeated multi-instance training and lives in one Availability Zone.
  • DynamoDB export to S3 needs PITR and uses no read capacity (full or incremental); RDS and Aurora snapshot export writes Parquet without querying the instance.
  • Intelligent-Tiering for unknown access with no retrieval fees; Glacier Deep Archive for the cheapest long-term retention; Object Lock compliance mode for immutability nobody can bypass; SSE-KMS with a customer managed key for key control.
  • A training job reads data from its own Region and has separate key settings for output artifacts (KmsKeyId) and instance volumes (VolumeKmsKeyId).
  • Hot Kinesis shards need a better partition key, not more shards; slow Lambda consumers need ParallelizationFactor; S3 503s need more prefixes and fewer, larger files.
  • Firehose for delivery without code, Kinesis Data Streams for replay and multiple consumers (enhanced fan-out for dedicated throughput), MSK for Kafka producers, Managed Flink for stateful windows.
  • Firehose format conversion needs JSON input and a Glue table; dynamic partitioning groups by record values; an object is written when either buffering hint is reached.
  • Parquet or ORC for column reads; built-in algorithm CSV has the target first and no header; Bedrock fine-tuning data is JSON Lines; Avro plus Glue Schema Registry enforces stream schema compatibility.
  • S3 Vectors for cheap infrequent queries, OpenSearch Serverless vector search for high-rate hybrid search, pgvector for vectors next to live relational data; index dimension must equal the embedding model's output.
How the exam tests this

Task 1.1 is tested with short scenarios about ML data pipelines: a training job that starts slowly or runs short of disk, an extraction that must not burden production, a stream or bucket that throttles or falls behind, data that must be cheaper to keep or must meet a compliance rule, a vector store that has to match an embedding model, unstructured data that must be ingested and made searchable, or features that must be served online and kept for training. Each scenario states one or two constraints, about code changes, operational effort, latency, cost, data location or freshness, and the right answer is the option that meets every stated constraint rather than the most powerful service. Options often differ in a single configuration detail, so knowing what each setting actually does, and which settings are fixed at creation, matters more than recognising service names. Some questions ask for more than one action.

Data Transformation, Feature Engineering and Pre-processing on AWS (MLA-C02)

  • Glue = serverless Spark (job bookmarks for incremental runs); DataBrew = visual no-code recipes; EMR = cluster control and Spot task nodes; Data Wrangler = ML-focused visual prep you export to Pipelines, Feature Store or a processing job.
  • Every feature group needs a record identifier and an event time; the online store serves the latest value in milliseconds, the offline store keeps every version in S3.
  • Build training sets with point-in-time joins on the event time feature, never on current values or write_time; online-store TtlDuration expires stale records while the offline store keeps history.
  • Lambda handles stateless per-record stream transforms and short tumbling windows; sliding or event-time windows with late data need Spark Structured Streaming with a watermark, or Flink.
  • Scale features for distance and gradient-based models, use a robust scaler when outliers must stay, and fit every transform on the training split only, then reuse it at inference.
  • Use log(1 + x) for right-skewed features with zeros, exponentiate predictions from a log target, and bin plus one-hot a non-monotonic feature for a linear model.
  • One-hot encode nominal categories, ordinal-encode ordered ones, and split timestamps into local-time components.
  • Documents and queries must share one embedding model and dimension; changing either means a new index and a full re-embed.
  • Chunking: no chunking for short self-contained files, hierarchical for precision plus context, semantic threshold lower for smaller chunks, custom Lambda for your own splitting; metadata via <file>.metadata.json plus a sync.
  • Redact the data itself (Comprehend, Glue sensitive data detection, DataBrew); Macie only discovers. Fine-tuning takes labeled pairs, continued pre-training unlabeled input text, distillation prompts and a teacher model.
How the exam tests this

Task 1.2 questions are scenarios about a transformation pipeline, feature, stream, vector index or training dataset that has to meet a constraint: who builds it, how much infrastructure the team will run, how fresh or historically accurate the features must be, how late data can arrive, what a particular model can learn, how documents must be retrieved, which data must stay private, or what kind of training data exists. The deciding constraint is often stated in passing, so the work is to identify it and pick the service, setting or technique it implies, and to recognise when a plausible alternative changes the data, the scale or the meaning in a way the scenario does not allow. Some questions ask for more than one action that together complete a change.

Data Quality Validation, Bias Metrics and Class Imbalance for ML (MLA-C02)

  • Glue Data Quality in an ETL job gives row-level results so bad rows can be quarantined in the same run; Data Catalog runs check data at rest, recommend rules and filter with partition predicates.
  • DQDL Boolean rules (IsComplete, IsUnique) allow zero failures; ratio rules (Completeness, Uniqueness) take a threshold; last(k) with k > 1 needs an aggregation such as min or avg.
  • A DataBrew ruleset is evaluated by a profile job, and the pass or fail lives in the Ruleset Validation Result event, not the job state.
  • SageMaker Ground Truth and SageMaker Clarify are closed to new customers; a new account labels with FM pre-labeling plus human audit and computes bias metrics itself.
  • Split by key for repeated entities, use an ordered split for time series, shuffle sorted exports and deduplicate before any split.
  • CI measures how many rows each facet has; DPL measures how often each facet gets the positive label; CDD conditions on a subgroup to expose Simpson's paradox.
  • Resample only the training split after splitting; when rows must stay as recorded, use scale_pos_weight (binary), XGBoost instance weights with csv_weights=1 (multi-class) or inverse-frequency loss weights.
  • Augment rare conditions with label-preserving transforms in training only, and measure on real held-out examples.
  • ApplyGuardrail screens text against an existing guardrail without invoking a model; Comprehend gives per-category toxicity scores; Rekognition DetectModerationLabels screens images.
  • Impute skewed numbers with the median and categories with the mode, fitted on the training split; treat sentinel codes as missing and keep informative extremes.
How the exam tests this

Task 1.3 questions are scenarios with one deciding constraint in the middle: who owns the data and whether they write code, whether bad rows may reach training, whether rows may be changed, where the data must stay, whether the account is new, or what range the values can physically take. Options are often close variants of one another that differ in a single detail, so read each rule, setting or event name in full rather than recognising a familiar service. Several questions describe a symptom rather than naming the problem (a test score far above production performance, a suspiciously low validation loss, a metric that misses its own outliers) and ask for the fix, sometimes as a Select TWO where both halves of a solution are needed. Bias questions give metric values or a fairness question and ask which metric or conclusion applies; know what each pre-training metric measures and which metrics need a trained model's predictions.

ML Model and Foundation Model (FM) Development

24% of the exam

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02)

  • Use the least custom approach that works: AI service, then FM or traditional ML, then a custom model; FMs are a poor fit for high-volume tabular or fixed-class prediction when labels exist.
  • Interpretability requirements decide the model: fixed published weights mean a linear model; SHAP values vary per prediction, importance rankings lack direction, monotonic constraints fix direction only.
  • Filter FMs on hard requirements (modality, context window, languages, Region) first, then compare candidates on your own prompts for quality, latency and cost.
  • The same embedding model must embed documents and queries; use a multimodal model on the images for visual search and a multilingual model for cross-language retrieval; fewer dimensions means less storage.
  • RAG handles changing facts, citations and access control; fine-tuning handles tone, layout and labeling behaviour; many systems need both.
  • Customization follows the data: labeled pairs for supervised fine-tuning, a reward function for reinforcement fine-tuning, a good teacher and prompts for distillation, unlabeled text for continued pre-training.
  • LoRA adapters cut training cost and let many variants share one base model on a SageMaker AI endpoint.
  • Hybrid search fixes exact identifiers, GraphRAG fixes multi-hop questions, structured data stores compute exact aggregates; a missing mandatory feature rules out a cheaper vector store.
  • Prompt caching, batch inference, Provisioned Throughput, prompt routing and cross-Region inference each solve one cost or capacity problem and carry one constraint.
  • Textract AnalyzeExpense for invoices, AnalyzeID for identity documents, Rekognition DetectModerationLabels for unsafe images, Transcribe for speech with PII redaction and custom vocabulary, Comprehend for text.
How the exam tests this

Task 2.1 questions are scenarios that describe a business problem and its data, then ask which approach, model, method or configuration fits. The deciding detail is usually one constraint in the middle of the story: no labeled data or no reward function, a regulator's explanation rule, exact identifiers in queries, questions that link several documents, a single-Region residency rule, long idle periods, a modified model architecture, export to edge devices, or a fixed cost ceiling at high volume. Distractors are workable-sounding near-twins that miss that one constraint (the same embedding model applied to the wrong input, a cheaper vector store without hybrid search, Spot training without checkpoints, a custom classifier instead of an entity recognizer), and the modern-sounding option, such as a foundation model for tabular prediction, is often wrong. Multi-response items commonly pair two tools that each fix a separate part of the problem.

Training, Hyperparameter Tuning and Fine-Tuning Models on AWS (MLA-C02)

  • Built-in algorithm choice follows labels and output: RCF for unlabeled anomaly scores, DeepAR for many related time series, XGBoost for tabular prediction, k-means only clusters.
  • Built-in CSV input means the label in the first column and no header; FastFile streams data and ShardedByS3Key gives each instance its own subset.
  • In script mode, hyperparameters arrive as command-line arguments, the model must be saved to SM_MODEL_DIR (/opt/ml/model), and requirements.txt in source_dir adds packages to the managed container.
  • Bayesian tuning learns from completed jobs, so high parallelism turns it into random search; use logarithmic scaling for ranges spanning orders of magnitude and a regex metric definition for custom scripts.
  • Warm start TransferLearning handles new data or a new algorithm version; IdenticalDataAndAlgorithm needs the same data and image.
  • Early stopping (AMT Auto, Hyperband, XGBoost early_stopping_rounds with a validation channel) saves wasted compute; use data parallelism when the model fits on one GPU and model parallelism or sharding when it does not.
  • Overfitting needs regularization or data, underfitting needs capacity or features, and catastrophic forgetting needs LoRA, mixed general data, a lower learning rate or fewer epochs.
  • Bagging reduces variance, boosting reduces bias, stacking learns how to weight diverse models, and cascades or Bedrock intelligent prompt routing cut cost by sending easy requests to smaller models.
  • Prompt engineering first; fine-tuning needs labeled pairs, continued pre-training uses unlabeled text, and distillation trains a smaller student from a teacher's responses to your prompts.
  • Changing the embedding model or dimension means re-embedding the corpus; hybrid search fixes exact identifiers, hierarchical chunking adds context, and overlap stops split sentences.
How the exam tests this

Task 2.2 questions are scenarios that describe a training job, tuning job, fine-tuning run or RAG pipeline that is slow, failing, inaccurate or too expensive, often with one constraint in the middle of the story: no labels, no custom container images, a fixed number of training jobs, no larger instance, no classifier to maintain, or an acceptable accuracy loss. Expect near-twin options that differ in one setting (Logarithmic versus ReverseLogarithmic scaling, TransferLearning versus IdenticalDataAndAlgorithm, FastFile versus a bigger volume, hybrid versus semantic search, data versus model parallelism), and multi-response items where two settings together solve two separate symptoms. Read the training and validation numbers carefully before deciding between overfitting and underfitting, and check what data the scenario actually has before choosing a customization method.

Evaluating ML and GenAI Models: Metrics, Drift, Explainability and RAG on AWS (MLA-C02)

  • Reproducible runs need the data version (MLflow dataset input with a digest) and the code version (Git commit tag), not only parameters and metrics; a training job reaches managed MLflow only with the tracking server ARN as tracking URI and the sagemaker-mlflow plugin.
  • In Bedrock Prompt Management, iterate on the draft, compare alternatives as variants, and ship immutable numbered versions by ARN; fix a prompt by creating a new version, never by editing or deleting the old one.
  • Baselines come from the training data. Input-distribution checks catch data drift; only ground-truth quality metrics, joined by inference ID when labels arrive late, catch concept drift.
  • SageMaker Model Monitor, Clarify and Debugger are closed to new customers: new accounts use data capture with Evidently, MLflow and SNS/CloudWatch for drift, the SHAP library and pandas or scikit-learn bias formulas for predictive models, fmeval or Bedrock evaluations for foundation models, and metric definitions, CloudWatch alarms and TensorBoard for training.
  • A shadow variant sees live traffic but never answers customers; it needs an instance-based real-time endpoint, data capture compares responses, and completing the test can deploy the shadow variant.
  • Local explanation means one prediction's SHAP values; global ranking means mean absolute SHAP; how the prediction changes as one feature moves is a partial dependence plot.
  • Divergence to NaN: lower the learning rate. Vanishing gradients: ReLU with He initialization. Exploding gradients: clip.
  • With rare positives, ignore accuracy and compare recall, precision or PR AUC; linear error cost means MAE, costly large misses mean RMSE.
  • BLEU for translation, ROUGE for summaries, BERTScore or embedding similarity when correct answers are paraphrased.
  • Bedrock evaluations: automatic for repeatable algorithmic scores, LLM-as-a-judge (with custom metrics and an independent, human-calibrated judge) for subjective quality, human jobs with your own team for expert or confidential review, retrieve-only RAG jobs to test retrieval and retrieve-and-generate jobs for faithfulness and citations.
How the exam tests this

Task 2.3 questions are scenarios that describe a model, prompt or RAG system that has to be compared, explained, debugged or monitored, usually with one constraint in the middle of the story: customers must see only production responses, the account is new (so tools closed to new customers are out), labels arrive weeks late, no human raters are available, scores must be repeatable, data is confidential, or the model must stay where it is hosted. Expect near-twin options that differ in one detail: a baseline from training data versus recent traffic, a retrieve-only versus retrieve-and-generate job, citation precision versus citation coverage, choice buttons versus Likert comparison, Stereotyping versus Harmfulness, local SHAP versus mean absolute SHAP versus a partial dependence plot, BLEU versus ROUGE versus BERTScore, MAE versus RMSE. Work out which component or failure mode the symptoms point to before choosing the metric or tool, and check whether the scenario needs an evaluation report or a runtime control.

Deployment and Orchestration of ML and AI Workflows

24% of the exam

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02)

  • Synchronous plus idle gaps plus acceptable cold start means serverless inference; large payloads or minutes of processing with scale-to-zero means asynchronous inference; a scheduled dataset means batch transform.
  • Asynchronous requests pass the payload as an S3 InputLocation, and success and error SNS topics in AsyncInferenceConfig report completion without polling.
  • Batch transform needs SplitType Line with MultiRecord batching for big files, parallelizes by file, and filters with InputFilter, JoinSource and OutputFilter in that order.
  • Inferentia2 is the low-cost inference accelerator only for Neuron-compatible models, Graviton needs an arm64 image, and GPU instances should be right-sized to the model's memory.
  • Multi-model endpoints share one container across many similar models (Triton on GPUs); multi-container Direct mode hosts different frameworks; serial pipelines chain containers; inference components give each model its own resources and scaling.
  • Large models fit only when weights plus KV cache fit the usable GPU memory: shard with tensor parallelism across every GPU, or quantize when the instance is fixed.
  • Bedrock on-demand bills per token, cross-Region inference profiles absorb peaks without commitment, Provisioned Throughput is a commitment, batch inference serves offline jobs, and Marketplace models run on endpoints you size.
  • Custom Model Import takes supported LLM architectures in Hugging Face safetensors format and serves them on demand, but not embedding models and not with batch inference.
  • Bedrock Agents act through action groups (Lambda or return of control, with optional user confirmation); AgentCore Runtime hosts framework agents; MCP connects agents to tools and A2A connects agents to agents.
  • Reranking fixes good chunks ranked too low, metadata filters fix wrong-scope answers, and query decomposition fixes one-sided multi-part answers.
How the exam tests this

Task 3.1 questions describe a model, foundation model, agent or RAG system and its traffic, latency, payload size, team skills and budget, then ask which deployment, configuration or setting fits. One constraint, usually stated mid-scenario, decides between options that would all work: the caller needs a synchronous answer, nothing may be paid while idle, the instance type is fixed, no capacity commitment, no extra hosted model, no Lambda functions, data must stay in one geography, or every service must use existing Kubernetes tooling. Expect near-twin options that differ in one component (serverless versus asynchronous inference, multi-model versus multi-container endpoints, Inferentia versus Trainium, geographic versus global inference profiles, MCP versus A2A, Retrieve versus RetrieveAndGenerate, a knowledge base reranking setting versus the Rerank API), memory arithmetic for GPU sizing, and multi-response items where two settings fix two separate symptoms.

Provisioning ML Resources: Endpoints, Scaling, Knowledge Bases and Agents (MLA-C02)

  • Only Bedrock Provisioned Throughput reserves model capacity for one workload; cross-Region inference profiles absorb bursts while staying on-demand, and batch inference is never interactive.
  • Serverless provisioned concurrency removes cold starts and can be scheduled with Application Auto Scaling, but its scalable target floors at 1, and serverless has no GPUs.
  • CloudFormation exports lock the producer while imported; Parameter Store dynamic references do not. In Step Functions, SageMaker AI .sync exists only for jobs, never .waitForTaskToken, so wait for endpoints with a Wait + DescribeEndpoint + Choice loop.
  • Inference containers answer GET /ping and POST /invocations on port 8080; training scripts must save the model to /opt/ml/model. Extend the AWS image FROM it when you need OS packages or have no internet.
  • VpcConfig belongs on the model. Only S3 and DynamoDB have gateway endpoints; runtime and control-plane APIs (sagemaker.runtime vs sagemaker.api, bedrock-runtime vs bedrock) are separate interface endpoints, and default SDK hostnames need private DNS.
  • Deploy with create_model, create_endpoint_config, create_endpoint; configurations are immutable, so swap versions with a new configuration and update_endpoint, and wait with the endpoint_in_service waiter on the endpoint name.
  • Scale short requests on invocations per instance, streaming LLMs on high-resolution concurrent requests, async on backlog, and variable-cost GPU work on GPUUtilization from /aws/sagemaker/Endpoints; scale from zero needs a step policy on NoCapacityInvocationFailures (or HasBacklogWithoutCapacity for async).
  • Inference components plus managed instance scaling give each model its own GPUs and scaling; training plans reserve GPU capacity, which quotas, Spot and Savings Plans do not.
  • Knowledge base vector fields must match the embedding model's dimensions; hybrid search needs OpenSearch Serverless, RDS/Aurora or MongoDB; Aurora needs the Data API and a Secrets Manager secret; metadata comes from .metadata.json sidecars.
  • AgentCore Runtime sessions are ephemeral ARM64 microVMs; replay exact turns from Memory events, keep distilled preferences with long-term strategies, and use Gateway for shared MCP tools and Identity for outbound OAuth tokens.
How the exam tests this

Task 3.2 questions describe an existing architecture (private networking, separately owned infrastructure stacks, a custom container, an endpoint that scales poorly, a knowledge base returning wrong chunks, an agent that loses state) and ask which configuration fixes it. One stated constraint usually decides between options that would all work, such as no outbound internet, no custom code, no idle cost or no long-term commitment, so read the constraints before the options. Expect options that look alike and differ in a single detail, such as the API, endpoint type, resource name or metric they use; knowing exactly what each service, endpoint and metric does is what separates them. Multiple-response items often combine two settings that each fix a separate symptom.

MLOps CI/CD Pipelines, Model Versioning and Rollback on AWS (MLA-C02)

  • CodePipeline releases, CodeBuild builds and tests, CodeConnections links external Git, and SageMaker Pipelines runs the ML workflow; CodeDeploy shifts traffic for inference on Lambda or ECS, while SageMaker AI endpoints use their own guardrails.
  • Docker builds in CodeBuild need privileged mode; ECR pushes need service-role permissions; a full-clone source needs codeconnections:UseConnection on the CodeBuild role too.
  • Connections created by CLI or CloudFormation stay PENDING until completed in the console; V2 push triggers filter by tag or by branch and file path, pull-request triggers by branch, file path and event type.
  • Blue/green needs a full second fleet (all at once, canary, linear); rolling updates replace capacity in batches for quota-limited fleets. Auto-rollback needs alarms in AutoRollbackConfiguration plus baking periods long enough to evaluate them.
  • Spot training keeps progress only with S3 checkpoints; MaxWaitTimeInSeconds covers waiting plus running and must be at least MaxRuntimeInSeconds.
  • Batch transform: SplitType splits, BatchStrategy packs, AssembleWith formats output, and InputFilter/JoinSource/OutputFilter reshape records in the same job.
  • Gate model quality with a ConditionStep and a FailStep; test deployed endpoints after the staging deploy; use a manual approval action for human sign-off.
  • Retrain from EventBridge schedules or from CloudWatch alarm state changes on Model Monitor metrics, never from every monitoring run; start releases from Model Package State Change events.
  • Knowledge base sync is incremental and includes deletions; a new embeddings model means a new knowledge base, a full ingestion and an ID switch.
  • Call prompts, custom models and agents through stable versioned identifiers: a versioned prompt ARN, a deployment or provisioned model ARN, a pinned AgentCore endpoint.
  • CodeDeploy releases need a deployment group, an AppSpec file, a canary, linear or all-at-once deployment configuration, and alarms on the deployment group for automatic rollback.
How the exam tests this

Task 3.3 tests whether you can configure, trigger and troubleshoot automated ML and generative AI release workflows on AWS: setting up and fixing CodePipeline, CodeBuild, CodeConnections, CodeCommit and CodeDeploy; orchestrating training, evaluation and batch inference with SageMaker Pipelines; placing unit tests, metric gates, integration tests and approvals in the right stage; choosing endpoint deployment strategies and rollback; triggering retraining and knowledge base refreshes from events; and versioning models, prompts, fine-tuned foundation models and agents so releases are repeatable and reversible. Questions are scenarios with a stated requirement or a failure symptom, and some are multiple response.

Operating, Monitoring, and Securing ML and AI Solutions

24% of the exam

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02)

  • Bedrock invocation logging is off by default; the Model Invocations view's per-request table needs it delivered to CloudWatch Logs, and bodies over 100 KB need an S3 large-data location.
  • TimeToFirstToken measures a slow start in streaming; InvocationLatency measures the whole response. Early failures appear only in Bedrock metrics without the ModelId dimension.
  • Every Model Monitor type needs data capture and a baseline; model quality also needs ground truth joined by InferenceId and a baseline built from predictions plus labels.
  • Data quality watches inputs, model quality watches accuracy, bias drift watches outcome gaps between groups, feature attribution drift watches the SHAP importance ranking.
  • Alarm on the drift metric, then let an EventBridge rule on the alarm state start the retraining pipeline; use anomaly detection bands for seasonal metrics and Logs anomaly detection for unpredictable log patterns.
  • Production variants expose users to the new model by weight; shadow variants never do. UpdateEndpointWeightsAndCapacities changes weights in place, TargetVariant pins a single request, InvokedProductionVariant attributes outcomes.
  • Bedrock judge-based and RAG evaluations can score your own logged responses without re-invoking the model; retrieve-only jobs isolate retrieval quality.
  • Handled tool errors are invisible to Lambda Errors; emit a custom metric. A stream without messageStop or with stopReason max_tokens is incomplete, whatever the HTTP status.
  • AgentCore online evaluation samples live sessions continuously; pick session, trace or tool-level evaluators to match the failure, and alarm on the Bedrock-AgentCore/Evaluations metrics.
How the exam tests this

Task 4.1 questions describe a model, foundation model, RAG system or agent in production and a symptom or requirement, then ask which monitor, metric, evaluation or configuration detects it. The deciding detail is usually a mid-scenario constraint: no ground-truth labels, labels arriving weeks later, no human reviewers, answers must be scored as delivered without re-invoking the model, no new AWS resources, users must never see the candidate's output, the agent's replies must not change, or as little custom code as possible. Expect near-twin options that differ in one component: data quality versus model quality versus bias drift versus feature attribution drift, production versus shadow variants, UpdateEndpoint versus UpdateEndpointWeightsAndCapacities, TargetVariant versus InvokedProductionVariant, InvocationLatency versus TimeToFirstToken, retrieve-only versus retrieve-and-generate metrics, online versus on-demand evaluation and session versus tool-level evaluators.

Optimising ML and GenAI Inference Cost and Performance on AWS: Instances, Savings Plans, Tokens and Vectors (MLA-C02)

  • Pick the instance family that supplies the bottleneck resource: c for CPU-bound, r for memory-heavy, GPU only when the model uses it, and check GPU memory (16, 24 or 48 GB) before LLM fit.
  • Inference Recommender default jobs shortlist instances; advanced load test jobs take your traffic phases, endpoint configurations and percentile-latency stopping conditions.
  • ModelLatency plus OverheadLatency is the time inside SageMaker AI; the rest of a slow request is found with X-Ray active tracing, whose sampling is decided at the entry service.
  • AgentCore traces need the ADOT SDK and CloudWatch Transaction Search enabled once per account; metrics appear without it, spans do not.
  • Dashboards: plot percentiles for latency, use cross-account observability for one central view, and SEARCH expressions so new resources appear automatically.
  • Tags reach Cost Explorer only after activation; per-team Bedrock attribution uses tagged application inference profiles; Budgets forecast alerts warn early and Budgets actions apply IAM policies or SCPs.
  • SageMaker AI commitments are SageMaker Savings Plans (not Compute Savings Plans or EC2 RIs), sized to the steady hourly baseline.
  • Bedrock Flex discounts latency-tolerant synchronous calls; batch discounts S3 jobs; Priority costs more; idle GPU endpoints are better replaced by per-token on-demand where the model is offered.
  • AgentCore Runtime bills CPU only while processing but memory until the session ends, so stop sessions and shorten the idle timeout.
  • Bedrock reserves input plus maximum output tokens against the quota; cache checkpoints go after the static prefix and must meet the model's minimum size; resync incrementally and shrink vectors to cut RAG cost.
How the exam tests this

Task 4.2 questions describe a model, foundation model, RAG pipeline or agent in production with a cost, latency or capacity symptom, usually backed by metric values, and ask which configuration, metric, tool or purchasing option fixes it. The deciding detail is often a mid-scenario constraint: no custom benchmarking harness, no copying of metrics, running work must not stop, requests must stay in one Region, calls must remain synchronous without S3 input files, keep cost close to today's, no fixed thresholds, or newly created resources must appear automatically. Expect near-twin options that differ in one component: ModelLatency versus OverheadLatency, MemoryUtilization versus GPUMemoryUtilization, default versus advanced Inference Recommender jobs, SageMaker versus Compute Savings Plans, baseline versus average commitment, Flex versus Priority versus batch, multi-model versus multi-container endpoints, Budgets forecast versus actual alerts, Cost Anomaly Detection versus CloudWatch anomaly detection, idle timeout versus maximum lifetime, and cache checkpoint placement. Some questions combine two facts, such as GPU memory per instance size, X-Ray sampling at the entry service, or how a commitment applies hour by hour.

Securing ML and AI Workloads on AWS: IAM, VPC Isolation, Auditing and Bedrock Guardrails (MLA-C02)

  • Any registry-side image scan runs after the push; a 'never stored' rule needs the Inspector SBOM Generator and Scan API inside the build, and 'keep re-checking old images' needs Inspector continuous enhanced scanning.
  • CodeGuru Reviewer is closed to new repository associations and CodeGuru Security is discontinued; Amazon Inspector code security provides SAST, SCA and IaC scanning for GitHub and GitLab.
  • A SageMaker AI job runs with its execution role; the caller needs iam:PassRole scoped to that role ARN with iam:PassedToService = sagemaker.amazonaws.com.
  • Cross-account or SSE-KMS artifacts need kms:Decrypt in both the key policy and the role's identity policy; cross-account model packages also need a model package group resource policy.
  • Prevent non-compliant jobs with explicit Deny statements on SageMaker AI condition keys, one per requirement; negated operators match a missing key, positive ones need IfExists or a Null check.
  • Network isolation blocks all outbound calls from the container, even inside a VPC, but input channels and output upload still work; a VPC without an S3 endpoint or NAT gives timeouts, not AccessDenied.
  • SageMaker InvokeEndpoint and Bedrock InvokeAgent, Retrieve and ApplyGuardrail are CloudTrail data events (off by default); CloudTrail never records prompts, which need Bedrock model invocation logging.
  • Applications on AWS should use IAM role credentials; short-term Bedrock API keys (at most 12 hours) suit key-only tools, and long-term keys tied to IAM users are for exploration only.
  • Match the guardrail policy to the need: denied topics for subjects, content filters for harm categories, input-tagged prompt attack filters for jailbreaks, sensitive information and regex filters to block or mask PII.
  • Enforce a guardrail with bedrock:GuardrailIdentifier using the versioned ARN and an explicit Deny, and keep that role away from RetrieveAndGenerate and InvokeAgent.
How the exam tests this

Task 4.3 questions describe an ML pipeline, SageMaker AI job, endpoint, Bedrock application or agent together with a security requirement or a failure, and ask which control, policy, setting or credential meets it. They test when each scanning option runs, how execution roles, iam:PassRole, KMS key policies and resource policies combine, how IAM condition keys and SCPs prevent non-compliant ML resources, what VPC configuration, network isolation and VPC endpoints each do, what CloudTrail, model invocation logging and AWS Config record, which credential type suits an application or tool, and which Bedrock Guardrails policy and enforcement mechanism fits a stated need, including risks specific to agents. Troubleshooting items present an error message or CloudTrail record and ask for the fix. Some items ask for two answers.

Facts fresh? Prove it.
Drill MLA-C02 practice questions with an explanation on every option.
Practice now

Where these MLA-C02 facts came from

MLA-C02 cheat sheet: your questions

Every key takeaway from all 12 MLA-C02 revision lessons — 120 facts — pulled onto one page and grouped by exam domain, with the exam pattern behind each topic. It is a compression of the notes, not a separate set of content, so it can never contradict them.

Source

Exam structure, domain weights and scoring on this page come from the official MLA-C02 exam guide.