SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
MLA-C02 · Domain 3

Deployment and Orchestration of ML and AI Workflows practice questions

Deployment and Orchestration of ML and AI Workflows is worth 24% of the MLA-C02 exam — the 2nd-heaviest of the 4 domains. Selecting deployment infrastructure for ML models, foundation models and agents, provisioning and scaling the resources behind them, and automating MLOps with CI/CD pipelines. 6 fully worked examples are further down this page, answers included.

Exam weight
24%
the 2nd-heaviest of the 4 domains
Questions
114
across 3 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Deployment and Orchestration of ML and AI Workflows questions, fully explained

Questions from the MLA-C02 bank mapped to domain 3, with the answer key and the reasoning behind every option. None of them repeat the examples on the main MLA-C02 practice page.

Question 1Deployment and Orchestration of ML and AI Workflows

A radiology startup runs a GPU-based segmentation model in SageMaker AI. Each request is a 300 MB scan, and inference takes 6 to 9 minutes per scan. Scans arrive unpredictably, sometimes none for a whole day. Clinicians need results within 20 minutes of upload and want a notification when a result is ready. The startup does not want to pay for GPU instances while no scans are waiting. Which deployment meets these requirements?

Choose one.

  • a
    A real-time endpoint on a GPU instance with auto scaling and a 15-minute model timeout

    Real-time endpoints accept small payloads and enforce a short invocation timeout, so 300 MB scans and multi-minute inference do not fit, and the endpoint keeps at least one GPU instance running.

  • b
    A serverless endpoint with 6,144 MB of memory and provisioned concurrency set to zero

    Serverless inference does not support GPUs, and its payload size and processing time limits are far below a 300 MB scan that takes minutes to segment.

  • c
    A batch transform job on a GPU instance that an hourly Amazon EventBridge schedule starts

    Batch transform handles large inputs and releases compute afterwards, but an hourly schedule plus job start-up time can leave a scan waiting well beyond the 20-minute requirement.

  • d
    An asynchronous endpoint on a GPU instance that scales to zero and publishes to Amazon SNS Correct

    Asynchronous inference takes payloads up to 1 GB from S3, allows processing up to one hour, can scale its instance count to zero when the queue is empty, and sends success or error notifications through Amazon SNS.

The concept

Asynchronous inference queues large, long-running requests and can scale to zero.

Why that’s the answer

The deciding facts combine: payload size (300 MB) and processing time (minutes) rule out real-time and serverless inference; the need to pay nothing while idle rules out an always-on endpoint; and the 20-minute deadline rules out a scheduled batch job. Asynchronous inference accepts requests via S3 up to 1 GB, processes for up to an hour, scales to zero instances when nothing is queued and notifies through SNS.

How to reason it out
  1. Compare the payload and processing time with each option's limits.
  2. Eliminate options that cannot use a GPU or that keep an instance running.
  3. Check the turnaround deadline against scheduled batch jobs.
  4. Choose asynchronous inference with scale-to-zero and SNS notifications.

Exam tip: Large payloads, minutes per request, near-real-time results, idle periods: asynchronous inference.

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.

Question 2Deployment and Orchestration of ML and AI Workflows

An ML engineer runs a SageMaker AI batch transform job on CSV files whose first column is a customer ID. The model was trained without that column and fails when it receives it. The downstream team needs each output line to contain only the customer ID and the prediction. Which data processing settings should the engineer use for the job?

Choose one.

  • a
    InputFilter "$[1:]", JoinSource "Input", OutputFilter "$[0,-1]" Correct

    The input filter drops the ID before the model sees the record, joining the source puts the full input record next to the prediction, and the output filter keeps only the first column (the ID) and the last (the prediction).

  • b
    InputFilter "$[1:]", JoinSource "Input", OutputFilter "$[-1]"

    Joining keeps the ID available, but an output filter of the last column alone writes only the prediction and drops the customer ID.

  • c
    InputFilter "$[0:]", JoinSource "Input", OutputFilter "$[0,-1]"

    This filter passes every column including the ID to the model, which was trained without it and fails.

  • d
    InputFilter "$[1:]", JoinSource "Input", OutputFilter "$[1:]"

    The joined record is filtered to drop its first column, which removes the customer ID the downstream team needs.

The concept

Batch transform can filter inputs, join them to predictions and filter the joined output.

Why that’s the answer

Batch transform's DataProcessing settings work in order: InputFilter selects what the model receives, JoinSource "Input" appends the prediction to the original input record, and OutputFilter selects what is written. Dropping column 0 on input, joining, then keeping column 0 and the last column yields "ID, prediction".

How to reason it out
  1. Remove the ID from what the model sees with the input filter.
  2. Join the original input so the ID is still available after inference.
  3. Use the output filter to keep only the ID and the appended prediction.

Exam tip: Filter in, join source, filter out: the batch transform recipe for carrying IDs through.

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.

Question 3Deployment and Orchestration of ML and AI Workflows

A retailer's SageMaker AI batch transform job scores one 40 GB CSV file of product records, one record per line. The job fails with errors saying the request payload is too large, even after the engineer raised the instance count from 1 to 4. The model scores one record at a time with no dependency between lines. How should the engineer configure the job?

Choose one.

  • a
    Keep the default split type, raise MaxPayloadInMB to its 100 MB maximum, and keep four instances

    With no splitting the whole 40 GB file is still sent as one request, which no payload limit can accommodate.

  • b
    Set AssembleWith to "Line" and AcceptType to "text/csv" while keeping four instances

    AssembleWith controls how results are joined in the output file; it does nothing to the size of the requests sent to the model.

  • c
    Set SplitType to "Line", BatchStrategy to "MultiRecord", and a MaxPayloadInMB that fits a mini-batch Correct

    Splitting on lines lets batch transform cut the file into records and pack them into mini-batches no larger than MaxPayloadInMB, so each request stays within the limit.

  • d
    Split the work across eight instances so that each instance receives a smaller share of the file

    Batch transform distributes whole files across instances, so a single large file still goes to one instance in one oversized request.

The concept

SplitType breaks large input files into records that batch transform sends in mini-batches.

Why that’s the answer

The payload error comes from sending the file unsplit. SplitType "Line" tells batch transform where record boundaries are; BatchStrategy "MultiRecord" packs as many records as fit under MaxPayloadInMB into each request. More instances do not help because files, not lines, are the unit distributed across instances.

How to reason it out
  1. Identify that the file is being sent to the container whole.
  2. Recall that instances share work per file, not per line.
  3. Split on the record delimiter and size the mini-batches with MaxPayloadInMB.

Exam tip: Big single file in batch transform: split on Line and cap the mini-batch payload.

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.

Question 4Deployment and Orchestration of ML and AI Workflows

A media company serves a transformer-based text classification model to millions of requests per hour around the clock. It wants the lowest cost per inference from accelerators that AWS designed specifically for inference, and its engineers are willing to compile the model with the AWS Neuron SDK. Which instance family should it deploy the SageMaker AI endpoint on?

Choose one.

  • a
    ml.trn1 instances with AWS Trainium accelerators

    Trainium is AWS's accelerator family built primarily for training; it is not the purpose-built inference choice.

  • b
    ml.p4d instances with NVIDIA A100 GPUs

    A100 GPUs run the model well but are high-end training-class hardware and cost far more per inference than an inference-specific accelerator.

  • c
    ml.c7g instances with AWS Graviton3 processors

    Graviton CPUs cut cost for CPU-bound models, but they are general-purpose processors rather than accelerators built for transformer inference at this volume.

  • d
    ml.inf2 instances with AWS Inferentia2 accelerators Correct

    Inferentia2 is AWS's accelerator designed for high-throughput, low-cost deep learning inference, and models are compiled for it with the Neuron SDK.

The concept

AWS Inferentia (ml.inf2) is the purpose-built accelerator for cost-efficient deep learning inference.

Why that’s the answer

The stem asks for accelerators designed specifically for inference and accepts Neuron compilation. That is Inferentia2. Trainium is the training counterpart that shares the Neuron SDK, A100 GPUs are costlier general accelerators, and Graviton is a CPU.

How to reason it out
  1. Note the requirement is inference-specific accelerators at the lowest cost per inference.
  2. Recall that the Neuron SDK targets both Inferentia and Trainium.
  3. Pick Inferentia2 (ml.inf2), the inference-focused one.

Exam tip: Inferentia = inference, Trainium = training; both use the Neuron SDK.

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.

Question 5Deployment and Orchestration of ML and AI Workflows

A bank hosts a LightGBM fraud model on a SageMaker AI real-time endpoint running ml.c5 instances. Traffic is steady, the model is CPU-bound, and latency is well within target. The bank wants to cut hosting cost without changing the model, and its serving image is currently built only for x86_64. What should an ML engineer do?

Choose one.

  • a
    Move the endpoint to ml.c7g Graviton instances and deploy the existing x86_64 image unchanged

    Graviton instances run arm64, so an image built only for x86_64 will not start on them.

  • b
    Move the endpoint to ml.inf2 instances and compile the LightGBM model with the Neuron SDK

    Inferentia accelerates deep learning models compiled with Neuron; a gradient-boosted tree model is not a fit, and the accelerator adds cost.

  • c
    Move the endpoint to ml.c7g Graviton instances and deploy an image rebuilt for arm64 Correct

    Graviton instances give better price-performance for CPU inference, and the serving image only needs rebuilding for the arm64 architecture; the model itself is unchanged.

  • d
    Move the endpoint to ml.g5 instances so that inference runs on an NVIDIA GPU instead

    A GPU instance costs more per hour, and a small CPU-bound tree model gains little from it, so cost would rise.

The concept

Graviton (arm64) instances lower the cost of CPU-bound inference but need an arm64 container image.

Why that’s the answer

For a CPU-bound tree model with steady traffic, moving to Graviton is the price-performance lever. The catch is architecture: the serving container must be built for arm64. Accelerators (Inferentia, GPUs) do not help a small tree model and raise cost.

How to reason it out
  1. Classify the model as CPU-bound and not deep learning.
  2. Pick the cheaper CPU family: Graviton.
  3. Check the container architecture: Graviton needs an arm64 image.

Exam tip: Graviton saves on CPU inference only with an arm64 image.

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.

Question 6Deployment and Orchestration of ML and AI Workflows

A company's platform team runs every production service on Amazon EKS and deploys with Helm charts through a GitOps controller, with shared logging, service mesh and cost tooling built around Kubernetes. A new PyTorch model needs GPU inference. Leadership requires the model service to use the same deployment and operations tooling as every other service. Where should the ML engineer deploy the model server?

Choose one.

  • a
    A SageMaker AI real-time endpoint on GPU instances, called from Kubernetes pods through a VPC endpoint

    A SageMaker endpoint serves GPU inference well, but it is deployed and operated outside the Kubernetes tooling, which breaks the stated requirement.

  • b
    A Kubernetes deployment on Amazon EKS that schedules the model server container onto GPU nodes Correct

    Running the model server as a Kubernetes workload on GPU nodes keeps Helm, GitOps, logging and cost tooling unchanged while providing GPU inference.

  • c
    An Amazon ECS service deployment on AWS Fargate that runs the model server container behind a load balancer

    Fargate tasks have no GPU support, and ECS is outside the Kubernetes tooling the company standardised on.

  • d
    An AWS Lambda function packaged as a container image that loads the PyTorch model at start-up

    Lambda functions run without GPUs, and Lambda is not managed through the cluster's Helm and GitOps tooling.

The concept

The deployment orchestrator follows the organisation's operating model as well as the model's needs.

Why that’s the answer

SageMaker AI hosting is usually the lowest-effort target, but the stated constraint is to reuse the Kubernetes tooling. Amazon EKS with GPU nodes meets both GPU and tooling requirements. Fargate and Lambda lack GPUs, and a SageMaker endpoint sits outside the GitOps flow.

How to reason it out
  1. List the hard requirements: GPU inference and the existing Kubernetes tooling.
  2. Eliminate targets without GPU support.
  3. Eliminate targets operated outside Kubernetes.
  4. Deploy on EKS GPU nodes.

Exam tip: When the organisation runs on Kubernetes and must stay there, EKS is the orchestrator.

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.

What MLA-C02 domain 3 tests, topic by topic

The official exam guide breaks Deployment and Orchestration of ML and AI Workflows into 3 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published MLA-C02 practice questions per topic in Deployment and Orchestration of ML and AI Workflows
TopicWhat it coversQuestions
Manage deployment infrastructure for ML and AI model typesExam guide task 3.1 (MLA-C02). Selecting compute environments and deployment targets; deployment orchestrators and multi-model or multi-container strategies; inference strategies (real-time, batch); foundation model deployment options; deploying models built outside AWS (SageMaker AI, Amazon Bedrock Custom Model Import); deploying and configuring agents, their tool integrations and agent communication protocols; FM hosting and resource allocation; RAG system configuration (retrieval strategies, reranking).38
Provision and configure resources for ML and AI workloads based on existing architecture and requirementsExam guide task 3.2 (MLA-C02). Balancing on-demand and provisioned resources; automating compute provisioning across stacks and orchestration services; building and maintaining containers for ML and AI; SageMaker AI endpoints in a VPC; deploying models programmatically (SageMaker AI Python SDK, AWS CLI, Boto3); auto scaling metrics; Amazon Bedrock knowledge bases (vector database configuration, document indexing, retrieval optimization); retrieval pipelines; agent state management; GPU resource scaling; agentic workflow infrastructure.38
Implement automated orchestration and continuous integration and continuous delivery (CI/CD) pipelines for MLOps and AI workloadsExam guide task 3.3 (MLA-C02). Automated deployment strategies and rollback; configuring and troubleshooting AWS CodeBuild, CodeCommit, CodeDeploy, CodePipeline and CodeConnections; training and inference jobs; automated testing in ML and AI pipelines; re-training mechanisms; model versioning for repeatability and audit (SageMaker Model Registry, MLflow on SageMaker AI); prompt management (Amazon Bedrock Prompt Management); agent deployment pipelines and agent versioning; AI model and prompt testing frameworks; FM deployment automation with fine-tuned model versioning; pipeline orchestration for RAG updates and knowledge base refresh.38
Total114

Revise Deployment and Orchestration of ML and AI Workflows before you drill it

Other MLA-C02 domains

Deployment and Orchestration of ML and AI Workflows: your questions

Deployment and Orchestration of ML and AI Workflows is domain 3 of the MLA-C02 exam guide and carries 24% of the scored content — the 2nd-heaviest of the 4 domains. On a 65-question paper that works out to roughly 16 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official MLA-C02 exam guide.