SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
MLA-C02 · Domain 1

Data Preparation for ML and AI practice questions

Data Preparation for ML and AI is worth 28% of the MLA-C02 exam — the heaviest of the 4 domains. Collecting and storing data for ML and AI, transforming it and engineering features, and validating data quality and managing bias — for traditional ML and foundation models alike. 6 fully worked examples are further down this page, answers included.

Exam weight
28%
the heaviest of the 4 domains
Questions
135
across 3 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Data Preparation for ML and AI questions, fully explained

Questions from the MLA-C02 bank mapped to domain 1, with the answer key and the reasoning behind every option. None of them repeat the examples on the main MLA-C02 practice page.

Question 1Data Preparation for ML and AI

A retailer keeps 2 TB of order history in an Amazon DynamoDB table that serves production traffic with provisioned capacity. Each week an ML engineer needs a full copy of the table in Amazon S3 to build a training dataset. The copy must not consume the table's read capacity or require custom code to page through items. What should the engineer do?

Choose one.

  • a
    Run a weekly AWS Glue ETL job that reads the table with parallel scans and writes to S3

    Parallel scans read the live table and consume its provisioned read capacity, competing with production traffic.

  • b
    Turn on DynamoDB Streams and have an AWS Lambda function write each item change to S3

    Streams capture only changes after they are enabled, so the existing items never reach S3, and the engineer would also have to write the Lambda code.

  • c
    Enable PITR and start a weekly DynamoDB export to S3 from EventBridge Scheduler Correct

    The export to S3 feature reads from the point-in-time recovery backup, not the live table, so it consumes no read capacity and needs no scanning code; EventBridge Scheduler can call the export API each week.

  • d
    Take a weekly on-demand backup and restore it to a new table for the training pipeline

    A restore creates another DynamoDB table, not files in S3; the pipeline would still have to scan that table to get the data out.

The concept

DynamoDB export to Amazon S3 builds full or incremental exports from point-in-time recovery data without touching table capacity.

Why that’s the answer

DynamoDB's export to S3 feature requires point-in-time recovery to be enabled and writes the table's data to an S3 prefix in DynamoDB JSON or Amazon Ion format. Because it reads from the continuous backup, it uses no read capacity and needs no paging code. The export has no built-in schedule, so a scheduler such as EventBridge Scheduler starts it each week. A Glue scan job also lands the data in S3, but it reads the live table and competes with production. Streams miss existing items, and a restored backup is another table rather than S3 data.

How to reason it out
  1. Note the constraints: full copy in S3, no read capacity, no paging code.
  2. Recall that table exports read from point-in-time recovery backups.
  3. Eliminate scans (consume capacity) and streams (miss existing items).
  4. Enable PITR and start an export to S3 each week from a scheduler.

Exam tip: Full DynamoDB copy in S3 without read capacity: PITR plus export to S3.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

Question 2Data Preparation for ML and AI

Data scientists want a monthly copy of several tables from an Amazon RDS for PostgreSQL production database in Amazon S3, stored in a columnar format they can query with Amazon Athena. The database administrator will not allow any extra queries against the production instance. Which approach meets these requirements?

Choose one.

  • a
    Run pg_dump from an EC2 instance each month and upload the CSV output to Amazon S3

    pg_dump queries the production instance, which the administrator forbids, and produces row-based text rather than a columnar format.

  • b
    Create an AWS DMS full-load task from the production instance to Amazon S3 as Parquet

    DMS can write Parquet, but a full load reads every table from the production instance, which adds the load the administrator ruled out.

  • c
    Export an automated or manual DB snapshot to Amazon S3, which writes the tables as Parquet Correct

    Snapshot export reads from the snapshot rather than the instance and writes the selected tables to S3 in Apache Parquet, a columnar format Athena reads directly.

  • d
    Create a read replica and have the data scientists query it from a SageMaker notebook

    A read replica keeps the data in PostgreSQL rather than putting columnar files in S3 for Athena, and replication still adds work on the source.

The concept

RDS and Aurora snapshot export to Amazon S3 writes table data as Apache Parquet without querying the running instance.

Why that’s the answer

Exporting a DB snapshot to S3 extracts data from the snapshot, so the production instance serves no extra queries, and the output is Parquet, which Athena can query efficiently. pg_dump and a DMS full load both read from the production instance. A read replica does not put columnar files in S3 at all.

How to reason it out
  1. Note the two requirements: columnar files in S3 and no queries on production.
  2. Recall that snapshot export works from a snapshot and outputs Parquet.
  3. Eliminate options that read from the live instance.
  4. Choose snapshot export to S3.

Exam tip: RDS data to S3 as Parquet without touching the instance: export a snapshot.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

Question 3Data Preparation for ML and AI

A company stores raw training datasets in Amazon S3. New datasets are read heavily for a few weeks; after that, some are reused for retraining at unpredictable times while others are never read again. Every dataset must stay readable within milliseconds, and the company wants lower storage cost without paying per-GB retrieval charges. Which storage class should the ML engineer use?

Choose one.

  • a
    S3 Intelligent-Tiering for the dataset objects Correct

    Intelligent-Tiering moves objects between access tiers based on observed access, keeps millisecond access in its frequent and infrequent tiers, and charges no retrieval fees.

  • b
    S3 Standard-IA through a lifecycle rule after 30 days

    Standard-IA keeps millisecond access but bills a per-GB retrieval charge, which the company wants to avoid for datasets reused at unpredictable times.

  • c
    S3 Glacier Flexible Retrieval through a lifecycle rule

    Glacier Flexible Retrieval needs a restore that takes minutes to hours, which breaks the millisecond access requirement.

  • d
    S3 One Zone-IA for the dataset objects after upload

    One Zone-IA lowers cost but still charges for retrievals and stores data in a single Availability Zone, so it fails the no-retrieval-fee requirement.

The concept

S3 Intelligent-Tiering suits data with unknown or changing access patterns: it tiers objects automatically with no retrieval fees.

Why that’s the answer

When access patterns are unpredictable, a fixed lifecycle transition guesses wrong for some objects. Intelligent-Tiering monitors access per object and moves it between tiers, keeping millisecond access in its default tiers, with a small monitoring charge but no retrieval fees. Standard-IA and One Zone-IA charge per-GB retrieval, and Glacier Flexible Retrieval is not readable in milliseconds.

How to reason it out
  1. Note the access pattern: heavy at first, then unpredictable.
  2. Note the constraints: millisecond reads and no retrieval charges.
  3. Eliminate IA classes (retrieval fees) and Glacier Flexible Retrieval (restore delay).
  4. Choose S3 Intelligent-Tiering.

Exam tip: Unpredictable access, millisecond reads, no retrieval fees: S3 Intelligent-Tiering.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

Question 4Data Preparation for ML and AI

A team trains a vision model for many epochs on 60 TB of images in Amazon S3 using several GPU instances at once. Profiling shows the GPUs are idle much of the time while waiting for data, and every new job repeats the same reads. The team wants a high-throughput shared file system that stays in sync with the bucket. Which actions should the ML engineer take? (Select TWO.)

Choose TWO.

  • a
    Create an FSx for Lustre file system with a data repository association to the S3 bucket Correct

    FSx for Lustre provides high-throughput, low-latency parallel file access and loads objects from the linked S3 bucket, so repeated epochs and jobs read from the file system.

  • b
    Run the training jobs in a VPC subnet in the same Availability Zone as the file system Correct

    SageMaker AI reaches FSx for Lustre only through a VPC, and the file system lives in one Availability Zone, so the job's subnet must map to that zone.

  • c
    Create an Amazon EFS file system in bursting throughput mode and copy the image data to it

    EFS is a shared file system, but bursting throughput scales with stored size and credits, so it is not the high-throughput choice for saturating many GPUs.

  • d
    Keep File input mode and attach larger EBS volumes of the same type to each training instance

    Larger volumes only hold more data; each job still copies the dataset from S3 before training, and the volumes are not shared between instances.

  • e
    Copy the images to an S3 Express One Zone directory bucket and keep File input mode

    Express One Zone lowers per-request latency, but in File mode every job still downloads the whole dataset to each instance before training, and a directory bucket is not a shared file system.

The concept

FSx for Lustre as a SageMaker AI training data source: S3-linked, high-throughput, VPC-attached, single-AZ.

Why that’s the answer

FSx for Lustre is built for throughput-bound training: it serves files in parallel at high throughput, links to an S3 bucket through a data repository association, and is mounted by the job at start-up regardless of dataset size. Two configuration facts come with it: the training job must run in a VPC, and because the file system is in one Availability Zone, the job's subnet should be in that zone. EFS in bursting mode is not the high-throughput choice, larger EBS volumes keep the per-job download, and a directory bucket read in File mode still copies the data for every job.

How to reason it out
  1. Identify the bottleneck: data loading from S3 on every epoch and job.
  2. Pick the storage built for high-throughput training reads that syncs with S3.
  3. Recall its networking requirement: a VPC subnet in the file system's Availability Zone.
  4. Eliminate options that keep a full download per job or lack throughput.

Exam tip: GPU-starved training on S3 data: FSx for Lustre linked to the bucket, job in a VPC subnet in its AZ.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

Question 5Data Preparation for ML and AI

A healthcare company stores clinical notes for model training in an Amazon S3 bucket. Regulators require that the company control the encryption keys and see an audit trail of each key's use. Training snapshots must be kept for seven years, and during that period no user, including administrators with full S3 permissions, can delete or overwrite them. Which actions meet these requirements? (Select TWO.)

Choose TWO.

  • a
    Turn on S3 Object Lock in compliance mode with a seven-year default retention period Correct

    Compliance mode prevents any user, including the root user, from deleting or overwriting a locked object version or shortening its retention.

  • b
    Set the bucket's default encryption to SSE-KMS with a customer managed key Correct

    A customer managed KMS key is under the company's control (key policy, rotation, disabling) and each use is recorded in AWS CloudTrail.

  • c
    Turn on S3 Object Lock in governance mode with a seven-year default retention period

    Governance mode lets users with the s3:BypassGovernanceRetention permission delete or shorten retention, so administrators could still remove snapshots.

  • d
    Set the bucket's default encryption to SSE-S3 with Amazon S3 managed keys

    SSE-S3 encrypts the data, but Amazon S3 owns and manages the keys, so the company neither controls them nor gets per-key usage records.

  • e
    Turn on S3 Versioning and add a lifecycle rule that expires noncurrent versions after seven years

    Versioning keeps older versions, but a user with permission can still delete any version, so it does not stop administrators from removing data.

The concept

Compliance storage on S3: Object Lock compliance vs governance mode, and SSE-KMS customer managed keys vs SSE-S3.

Why that’s the answer

Two separate requirements need two controls. For immutability that even administrators cannot bypass, Object Lock must use compliance mode; governance mode can be bypassed with a specific permission. For key control and auditing, SSE-KMS with a customer managed key gives the company the key policy and logs every encrypt and decrypt call in CloudTrail; SSE-S3 keys are managed by S3. Versioning alone does not stop deletion.

How to reason it out
  1. Split the requirements: key control with audit, and undeletable retention.
  2. For keys, choose a customer managed KMS key over S3 managed keys.
  3. For retention, compare Object Lock modes: only compliance mode binds administrators.
  4. Reject versioning, which does not block deletes.

Exam tip: No one can delete it: Object Lock compliance mode. Company controls and audits keys: SSE-KMS customer managed key.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

Question 6Data Preparation for ML and AI

An IoT platform sends sensor readings to an Amazon Kinesis data stream with eight provisioned shards, using the device's region code as the partition key. There are three region codes. Producers receive ProvisionedThroughputExceededException errors although total traffic is well below the stream's aggregate write capacity, and CloudWatch shows one shard near its limit. What should the ML engineer change?

Choose one.

  • a
    Double the shard count with UpdateShardCount and keep the region code partition key

    Each partition key value still hashes to a single shard, so adding shards leaves the same three keys concentrated on the same hot shards.

  • b
    Switch the stream to on-demand capacity mode and keep the region code partition key

    On-demand mode adds shards as traffic grows, but one key value still maps to one shard, so the hot shard stays hot.

  • c
    Register the consumers for enhanced fan-out so the busy shard gets more throughput

    Enhanced fan-out raises read throughput for consumers; the errors here come from producers exceeding a shard's write limit.

  • d
    Use the device ID as the partition key so the records spread across the shards Correct

    With only three partition key values, at most three shards receive data; a high-cardinality key such as the device ID spreads writes across all eight shards.

The concept

Kinesis Data Streams partition keys: a low-cardinality key creates hot shards that more shards cannot fix.

Why that’s the answer

Kinesis hashes each partition key to one shard, and each shard has its own write limit. With three region codes, writes land on at most three shards, so one shard throttles while the stream as a whole is mostly idle. Adding shards, by count or through on-demand mode, does not change where a given key lands. Enhanced fan-out is a consumer-side feature. Spreading writes with a high-cardinality key fixes the imbalance.

How to reason it out
  1. Note the symptom: throttling on one shard while total traffic is low.
  2. Look at the partition key: three distinct values.
  3. Recall that one key value always maps to one shard.
  4. Choose a high-cardinality partition key.

Exam tip: Throttling on one shard with spare stream capacity: fix the partition key, not the shard count.

Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.

What MLA-C02 domain 1 tests, topic by topic

The official exam guide breaks Data Preparation for ML and AI into 3 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published MLA-C02 practice questions per topic in Data Preparation for ML and AI
TopicWhat it coversQuestions
Collect and store dataExam guide task 1.1 (MLA-C02). Extracting data from sources (Amazon S3, EBS, EFS, RDS, DynamoDB, OpenSearch Service); storage decisions by cost, performance, data structure and compliance; troubleshooting ingestion and storage capacity/scalability issues; streaming ingestion (Amazon Kinesis, Apache Flink, Apache Kafka); data formats by access pattern (Apache Parquet, JSON, CSV, ORC); merging sources (code, AWS Glue, Apache Spark); scalable vector databases for AI (OpenSearch Service, Amazon RDS with pgvector, Amazon S3); ingesting and storing text, image and audio data; ingesting into SageMaker Feature Store.45
Perform data transformation, feature engineering, and pre-processingExam guide task 1.2 (MLA-C02). Transforming data with AWS tools (AWS Glue, Glue DataBrew, Spark on Amazon EMR, SageMaker Data Wrangler); creating and managing features (SageMaker Feature Store); transforming streaming data (AWS Lambda, Spark); feature engineering (scaling, standardization, feature splitting, binning, log transformation, normalization); embedding models for text and image data; advanced text pre-processing (tokenization, domain-specific augmentation); preparing documents for Retrieval Augmented Generation (chunking strategies, metadata extraction); masking, redacting and anonymizing data; preparing data for FM fine-tuning, continuous pre-training and model distillation.45
Validate data quality and manage biasExam guide task 1.3 (MLA-C02). Validating data quality (Glue DataBrew, AWS Glue Data Quality); labeling and annotating data; identifying and mitigating bias in data (dataset splitting, shuffling, augmentation); applying bias metrics across numeric, text and image assets; resolving class imbalance; validating AI training-data integrity (prompt-response pair validation, content safety screening); cleaning data (outlier detection, imputing missing values, deduplication).45
Total135

Revise Data Preparation for ML and AI before you drill it

Other MLA-C02 domains

Data Preparation for ML and AI: your questions

Data Preparation for ML and AI is domain 1 of the MLA-C02 exam guide and carries 28% of the scored content — the heaviest of the 4 domains. On a 65-question paper that works out to roughly 18 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official MLA-C02 exam guide.