SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
MLA-C02 practice

Free MLA-C02 practice questions

Drill exam-realistic AWS Certified Machine Learning Engineer - Associate questions by domain, with an explanation on every option — not just the right one. 5 fully worked examples are further down this page, answers included.

Question bank
477
across 4 domains
Free, no account
5/day
sign up free to remove the cap
Real exam
65 Qs
130 min · pass 720 / 1000
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across all domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

5 sample MLA-C02 questions, fully explained

Real questions from the MLA-C02 bank, with the answer key and the reasoning behind every option. Read them, then go back up the page and try the rest.

Question 1Data Preparation for ML and AI

An ML engineer trains a PyTorch model on 3 TB of TFRecord shards stored under one Amazon S3 prefix. Each SageMaker AI training job uses the default input mode and spends about 40 minutes copying data before the first epoch begins. The training script opens the shards through ordinary file paths, and the team does not want to change the script or provision new storage. What should the engineer do?

Choose one.

  • a
    Set the S3 channel's input mode to FastFile and keep the existing S3 prefix Correct

    FastFile mode presents the S3 objects through a POSIX file interface and streams them on demand, so training starts without the up-front copy and the script's file reads keep working unchanged.

  • b
    Set the S3 channel's input mode to Pipe and keep the existing S3 prefix

    Pipe mode streams data into named FIFO pipes that the script has to read as a stream, so the file-path code would have to be rewritten.

  • c
    Keep File mode on the S3 channel and raise the training volume size for the dataset

    A larger volume avoids running out of disk but File mode still downloads the full 3 TB before training starts, so the 40-minute delay stays.

  • d
    Create an FSx for Lustre file system linked to the bucket and mount it as the channel

    FSx for Lustre also avoids the copy, but it is new storage to provision and needs the job attached to a VPC subnet, which the team ruled out.

The concept

SageMaker AI training input modes: File (copy first), FastFile (stream with file semantics) and Pipe (stream through FIFO pipes).

Why that’s the answer

File mode, the default, copies the whole dataset to the instance's storage before the script runs, which is where the 40 minutes go. FastFile mode keeps POSIX file access but streams objects from S3 as the script reads them, so the job starts almost immediately and needs no code change. Pipe mode also streams but changes the interface to named pipes, and FSx for Lustre works but adds a file system and VPC setup the team does not want.

How to reason it out
  1. Identify the cause: File mode downloads the full dataset before training starts.
  2. Note the constraints: no script change and no new storage service.
  3. Rule out Pipe mode (different read interface) and FSx for Lustre (new infrastructure).
  4. Choose FastFile mode on the same S3 prefix.

Exam tip: Slow start in File mode with a file-based script and S3 data: switch to FastFile.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

Question 2ML Model and Foundation Model (FM) Development

A legal-tech company is building an assistant on Amazon Bedrock that answers questions about individual contracts. Each contract can run to 400 pages, and the product team wants the whole contract sent to the model in a single request so that clauses that refer to each other are read together. Which model attribute should the ML engineer check first when shortlisting foundation models?

Choose one.

  • a
    The maximum number of output tokens the model returns per response

    Output length limits how long an answer can be. The requirement here is about how much text goes in, which the output limit does not govern.

  • b
    The number of dimensions in the model's embeddings

    Embedding dimensions matter for vector search with an embedding model. The assistant sends the contract to a text generation model, so this attribute does not decide whether the document fits.

  • c
    The on-demand price for each thousand output tokens

    Price matters when comparing models that already meet the requirement. A cheaper model whose context window cannot hold the contract fails the stated need.

  • d
    The maximum context window, measured in input tokens Correct

    A 400-page contract has to fit in the prompt in one request, so the context window is the first hard filter; a model whose window is too small is out regardless of quality or price.

The concept

Shortlisting a foundation model starts with hard task requirements such as input size, modality and language; price and quality are compared among models that pass them.

Why that’s the answer

The deciding requirement is that a document of hundreds of pages goes into one request. The context window (the input tokens a model accepts) is therefore a pass/fail filter. Output token limits, embedding dimensions and price are real selection criteria, but each is compared only among models whose context window can hold the contract.

How to reason it out
  1. Find the hard requirement: the full contract in one request.
  2. Translate it into a model attribute: input tokens accepted, the context window.
  3. Filter the model catalog on that attribute first.
  4. Compare price, latency and quality among the models that remain.

Exam tip: Hard requirements (context window, modality, language) filter the list; cost and quality rank what is left.

Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02) — the lesson that teaches this.

Question 3Deployment and Orchestration of ML and AI Workflows

An insurer retrains a claims-risk model in SageMaker AI every month. Each night, about 20 million new claim records land in Amazon S3 as CSV files, and every record must be scored before analysts arrive at 8 AM. No application calls the model during the day. The team wants to avoid paying for compute while nothing is being scored. Which inference option should the team use?

Choose one.

  • a
    Host the model on a real-time endpoint with a scheduled scaling action

    A real-time endpoint can be invoked in a loop, but it keeps at least one instance running and adds per-request overhead for a job that is purely offline.

  • b
    Run a SageMaker AI batch transform job over the nightly S3 prefix Correct

    Batch transform provisions instances for the job, scores the whole S3 dataset offline, writes results back to S3 and releases the compute when it finishes, so nothing is billed between runs.

  • c
    Host the model on a serverless endpoint and invoke it once per claim record

    Serverless scales to zero, but 20 million individual synchronous calls are slow and expensive compared with one offline job reading S3 directly.

  • d
    Host the model on an asynchronous endpoint that scales in to zero instances

    Asynchronous inference suits individual large or long-running requests arriving over time; queuing millions of tiny records through it is the wrong fit for a bulk nightly dataset.

The concept

Batch transform is SageMaker AI's offline inference option for scoring whole datasets in S3.

Why that’s the answer

The workload is a large dataset scored on a schedule with no interactive caller. Batch transform reads the input from S3, spins up the instances for the job only, writes predictions to S3 and stops. The endpoint-based options all exist to answer requests as they arrive; they can be bent to this job but cost more and add request plumbing.

How to reason it out
  1. Notice that no application needs an answer in real time.
  2. Notice the input is a complete dataset already in S3.
  3. Match offline scoring of a stored dataset with no idle cost to batch transform.

Exam tip: Whole dataset in S3, no live caller: batch transform, not an endpoint.

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.

Question 4Operating, Monitoring, and Securing ML and AI Solutions

A company runs a customer-support assistant on Amazon Bedrock. The operations lead wants a prebuilt view of invocation counts, token usage by model, latency percentiles and throttles. The lead also wants to open any single request and read the prompt and the model's response. The team does not want to build or maintain custom dashboards. What should the ML engineer do?

Choose one.

  • a
    Enable Bedrock model invocation logging to a CloudWatch Logs log group, and use the Model Invocations view in CloudWatch generative AI observability. Correct

    The Model Invocations view is prebuilt (invocations, tokens, latency, throttles, errors) and, once invocation logs go to CloudWatch Logs, lists each request with its input and output.

  • b
    Enable Bedrock model invocation logging with an Amazon S3 bucket as its sole destination, and use the Model Invocations view in CloudWatch generative AI observability.

    The metric widgets would work, but the request-level table reads invocation logs from CloudWatch Logs. With an S3-only destination the prompts and responses are not available in that view.

  • c
    Turn on AWS CloudTrail data events for Amazon Bedrock, and build a CloudWatch dashboard from the trail's log group with Logs Insights widgets.

    CloudTrail records who called which API, not the prompt and response bodies, and this option means building and maintaining a custom dashboard.

  • d
    Add the AWS/Bedrock InvocationLatency and token count metrics to a new CloudWatch dashboard, and use metric math to plot P90 and P99 latency.

    This is the custom dashboard the team wants to avoid, and metrics alone never show an individual request's prompt or response.

The concept

CloudWatch generative AI observability provides a prebuilt Model Invocations view for Amazon Bedrock; request-level inputs and outputs appear only when model invocation logging is sent to CloudWatch Logs.

Why that’s the answer

Two needs decide this: a prebuilt metrics view and per-request prompt/response detail. The Model Invocations view in CloudWatch generative AI observability covers the metrics out of the box, and its invocation table is populated from Bedrock model invocation logs delivered to CloudWatch Logs. Logging only to S3 leaves that table empty, CloudTrail never records payloads, and a hand-built metric dashboard is what the team wants to avoid.

How to reason it out
  1. List the requirements: prebuilt metrics, per-request content, no custom dashboards.
  2. Recognise the prebuilt GenAI view in CloudWatch for Bedrock invocations.
  3. Remember that request content comes from model invocation logging, which is off by default.
  4. Check the destination: the view reads the logs from CloudWatch Logs, not from S3.

Exam tip: Prebuilt Bedrock invocation dashboards come from CloudWatch generative AI observability; prompts and responses need invocation logging to CloudWatch Logs.

Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.

Question 5Data Preparation for ML and AI

A data-parallel training job runs on eight ml.p4d instances and reads 24,000 image archives from one Amazon S3 prefix in File mode. Every instance runs out of local storage during the download. The training script opens the archives as ordinary files from a local path, does no sharding of its own and assumes that each instance receives a different part of the dataset. The script must not be changed. Which channel configuration meets these requirements?

Choose one.

  • a
    File input mode with the S3 data distribution type set to FullyReplicated

    FullyReplicated is the default that caused the problem: every instance downloads the whole dataset, and every instance would train on the same data.

  • b
    Pipe input mode with the S3 data distribution type set to ShardedByS3Key

    Sharding is right, but Pipe mode streams the data into a FIFO pipe instead of files on disk, so the script would have to be rewritten to read from the pipe.

  • c
    File input mode with the S3 data distribution type set to ShardedByS3Key Correct

    ShardedByS3Key gives each instance a distinct eighth of the objects, which matches the script's assumption and cuts each instance's download to an eighth of the dataset, while File mode keeps the local file paths the script already reads.

  • d
    Pipe input mode with the S3 data distribution type set to FullyReplicated

    Pipe mode removes the disk pressure only by streaming into a FIFO pipe, which needs a script change, and replication still hands every instance the whole dataset.

The concept

Two independent channel settings: the input mode (File, FastFile, Pipe) decides how data reaches the container; the S3 data distribution type (FullyReplicated, ShardedByS3Key) decides which objects each instance gets.

Why that’s the answer

Two facts must be combined. First, the default FullyReplicated distribution gives each instance the entire dataset, which is why every instance runs out of space and why the workers would train on identical data; ShardedByS3Key splits the objects across the instances. Second, Pipe mode delivers data through a FIFO pipe rather than as files, so it breaks the unchanged script that opens files from a local path. File mode with ShardedByS3Key fixes both the storage and the data split without touching the script.

How to reason it out
  1. Separate the two settings: input mode and data distribution type.
  2. Distribution: the script expects distinct data per worker, so choose ShardedByS3Key.
  3. Input mode: the script reads local files and must not change, so rule out Pipe mode.
  4. Combine them: File mode with ShardedByS3Key.

Exam tip: Distinct data per worker is ShardedByS3Key; reading ordinary files without code changes rules out Pipe mode.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

The MLA-C02 question bank, by domain

The bank is built to the exam's own weighting, so the practice you get reflects the marks that are actually on offer — not whichever domain was easiest to write questions for.

Published MLA-C02 practice questions per exam domain
DomainExam weightTopicsQuestions
Data Preparation for ML and AI28%3135
ML Model and Foundation Model (FM) Development24%3114
Deployment and Orchestration of ML and AI Workflows24%3114
Operating, Monitoring, and Securing ML and AI Solutions24%3114
Total100%12477

Practise MLA-C02 one domain at a time

How MLA-C02 questions are worded

Most MLA-C02 questions are not asking whether you can recall a definition. They describe a situation and ask which option satisfies it — so the skill being tested is reading the requirement precisely and eliminating options that fail it.

Single-response vs multiple-response

A single-response question has exactly one right answer. A multiple-response question tells you how many to pick ("Choose TWO") and there is no partial credit — getting one of the two right scores nothing. Read that instruction before you read the options.

Read the last line first

The final sentence is the actual question; everything before it is scenario. Read it first, then read the scenario knowing what you are looking for. It stops you from building an answer in your head that the question never asked for.

Hunt for the qualifier

Most scenarios turn on one word — MOST cost-effective, LEAST operational overhead, with the LEAST latency, without changing application code. Two options are frequently both technically correct, and the qualifier is the only thing separating them.

Eliminate, then choose

Distractors are almost always real AWS services doing a real job — just not this job. Rule out the ones that break a stated constraint before you compare what's left. On a question you truly don't know, eliminating two options turns a guess into a coin flip.

Practise MLA-C02 from your browser

The SaveMyCert Chrome extension opens a daily drill of 5 MLA-C02 questions in the side panel, with the full explanation behind every answer. No account needed.

Get the Chrome extension (opens in a new tab)

Beyond MLA-C02 practice

MLA-C02 practice questions: your questions

Yes. You can answer 5 questions a day with no account at all, and a free account raises that to 30 a day across the full 477-question MLA-C02 bank, with an explanation on every option. Pro removes the daily cap entirely.

Source

Exam structure, domain weights and scoring on this page come from the official MLA-C02 exam guide.