SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
DEA-C01 · Domain 1

Data Ingestion and Transformation practice questions

Data Ingestion and Transformation is worth 34% of the DEA-C01 exam — the heaviest of the 4 domains. Ingesting streaming and batch data, transforming and processing it, orchestrating ETL pipelines, and applying programming concepts. Official weighting 34%. 6 fully worked examples are further down this page, answers included.

Exam weight
34%
the heaviest of the 4 domains
Questions
80
across 4 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Data Ingestion and Transformation questions, fully explained

Questions from the DEA-C01 bank mapped to domain 1, with the answer key and the reasoning behind every option. None of them repeat the examples on the main DEA-C01 practice page.

Question 1Data Ingestion and Transformation

A company runs an on-premises Oracle database that supports a live ordering system. The data engineering team must continuously replicate new inserts, updates, and deletes into an Amazon S3 data lake with minimal impact on the source database. Which approach should the team use?

Choose one.

  • a
    Schedule a nightly AWS Glue JDBC job that reloads all Oracle tables into S3

    A nightly full reload is batch, not continuous, misses intra-day changes, cannot easily represent deletes, and repeatedly running full-table scans places heavy query load on the production database.

  • b
    Use Amazon AppFlow to sync the Oracle tables to S3

    AppFlow is built for SaaS application sources such as Salesforce and ServiceNow; a self-managed on-premises Oracle database is not the kind of source AppFlow is designed to connect to.

  • c
    Enable Amazon DynamoDB Streams on the ordering tables

    DynamoDB Streams captures item-level changes only from DynamoDB tables; it cannot be attached to an Oracle database.

  • d
    Use AWS DMS with a full-load-and-CDC task targeting Amazon S3 Correct

    DMS change data capture reads the database's transaction logs to replicate ongoing inserts, updates, and deletes continuously with low impact on the source, and S3 is a supported DMS target for building a data lake.

The concept

AWS DMS change data capture (CDC) reads a relational database's transaction logs to stream ongoing changes to a target such as S3, Kinesis, or Redshift — the standard pattern for continuously replicating an operational database into a data lake.

Why that’s the answer

A DMS full-load-and-CDC task first copies existing data, then continuously applies inserts, updates, and deletes by reading redo logs, which keeps the load on the live Oracle system minimal. Nightly Glue reloads are high-impact batch jobs that miss changes between runs, AppFlow targets SaaS APIs rather than on-premises databases, and DynamoDB Streams only exists for DynamoDB.

How to reason it out
  1. Classify the requirement: continuous replication of row-level changes (CDC), not periodic snapshots.
  2. Choose AWS DMS, which supports Oracle as a source via transaction-log-based CDC.
  3. Configure a full-load-and-CDC task so historical data is copied once and ongoing changes flow afterward.
  4. Set Amazon S3 as the DMS target to land change records in the data lake.

Exam tip: For continuous, low-impact replication of a relational database's changes into AWS, use DMS with CDC.

Data Ingestion on AWS: Kinesis, MSK, DMS, Glue, and Batch Patterns — the lesson that teaches this.

Question 2Data Ingestion and Transformation

An application stores customer profiles in an Amazon DynamoDB table. The data engineering team must trigger processing logic whenever an item in the table is created or modified, without changing the application code. Which solution meets this requirement?

Choose one.

  • a
    Modify the application to publish a message to Amazon SNS after every write

    Dual-writing from the application would work but explicitly violates the requirement to avoid changing application code, and it risks missed events if the app writes to DynamoDB but fails to publish.

  • b
    Run an AWS Glue job every five minutes that scans the table for changed items

    Periodic full-table scans are expensive in read capacity, add up to five minutes of latency, and cannot reliably detect which items changed without extra timestamp bookkeeping.

  • c
    Create an Amazon S3 Event Notification on the table

    S3 Event Notifications only fire for object events in S3 buckets; they cannot be attached to a DynamoDB table.

  • d
    Enable DynamoDB Streams on the table and configure an AWS Lambda event source mapping Correct

    DynamoDB Streams emits an ordered record of item-level changes automatically, and a Lambda event source mapping polls the stream and invokes the function for each batch — no application changes are required.

The concept

DynamoDB Streams provides a time-ordered change log of item-level modifications, and Lambda event source mappings consume it automatically — the native change-driven ingestion pattern for DynamoDB.

Why that’s the answer

Enabling the stream requires only a table setting, and the Lambda event source mapping handles polling, batching, and retries, so processing reacts to every create and update with zero application changes. Dual-writes require code changes and can drift out of sync, scheduled scans are costly and slow, and S3 notifications do not apply to DynamoDB.

How to reason it out
  1. Enable DynamoDB Streams on the table and pick a view type such as NEW_AND_OLD_IMAGES so processors see the changed data.
  2. Create a Lambda function containing the processing logic.
  3. Add an event source mapping from the stream to the function; Lambda polls shards and invokes the function with batches of change records.
  4. Configure error handling on the mapping (retries, on-failure destination) so bad records do not block the shard.

Exam tip: To react to DynamoDB item changes without touching application code, pair DynamoDB Streams with a Lambda event source mapping.

Data Ingestion on AWS: Kinesis, MSK, DMS, Glue, and Batch Patterns — the lesson that teaches this.

Question 3Data Ingestion and Transformation

A data engineer must load 500 GB of gzip-compressed CSV files from Amazon S3 into an Amazon Redshift cluster as fast as possible. The data is currently stored as a single large file. What should the engineer do to maximize load performance?

Choose one.

  • a
    Load the single file with multiple concurrent INSERT statements

    Row-by-row or multi-row INSERTs are dramatically slower than COPY for bulk loads because they do not use Redshift's parallel ingestion path and generate far more per-statement overhead.

  • b
    Run one COPY command per file sequentially against the same table

    Issuing sequential COPY commands serializes the load; a single COPY with a manifest or common prefix lets Redshift parallelize across all files at once.

  • c
    Decompress the file first, because Redshift COPY cannot read gzip files

    COPY reads gzip-compressed files directly with the GZIP option; decompressing first only increases the data volume transferred and adds an unnecessary step.

  • d
    Split the data into multiple similarly sized files and load them with a single COPY command Correct

    Redshift's COPY command loads files in parallel across the cluster's slices, so splitting the data into multiple similarly sized files (ideally a multiple of the slice count) lets every slice work simultaneously instead of one slice processing a single huge file.

The concept

Amazon Redshift COPY achieves its speed through massively parallel loading: each slice in the cluster can ingest a file simultaneously, so file layout in S3 directly determines load parallelism.

Why that’s the answer

One large file forces effectively serial ingestion, while many similarly sized files loaded by a single COPY command spread the work across all slices. Concurrent INSERTs bypass the optimized bulk path, sequential COPY commands serialize the work, and COPY natively supports gzip so decompression is unnecessary.

How to reason it out
  1. Split the source data into multiple files of roughly equal size, ideally a multiple of the total slice count of the cluster.
  2. Keep the files compressed (gzip is supported) to reduce network transfer from S3.
  3. Issue a single COPY command with a common S3 prefix or a manifest file so Redshift loads all files in parallel.
  4. Avoid single-row INSERT patterns for bulk data — reserve them for small trickle loads.

Exam tip: For fast Redshift ingestion, split input into many equal files and load them with one parallel COPY command.

Data Ingestion on AWS: Kinesis, MSK, DMS, Glue, and Batch Patterns — the lesson that teaches this.

Question 4Data Ingestion and Transformation

A marketing team needs to ingest campaign data from Salesforce into Amazon S3 every night. The data engineering team wants a solution that requires no custom code and no infrastructure to manage. Which service should they use?

Choose one.

  • a
    An AWS Lambda function that calls the Salesforce REST API using a custom connector library

    This works but requires writing and maintaining custom API integration code, handling pagination, auth token refresh, and API limits — exactly the undifferentiated work AppFlow eliminates.

  • b
    AWS DMS with Salesforce configured as a source endpoint

    DMS sources are databases (and some data warehouses), not SaaS application APIs like Salesforce, so this configuration is not supported.

  • c
    Amazon AppFlow with a scheduled flow from Salesforce to Amazon S3 Correct

    AppFlow is a fully managed integration service with a native Salesforce connector; a scheduled flow moves SaaS data to S3 on a nightly cadence with configuration only — no code, no servers.

  • d
    An Amazon EMR cluster running a nightly Spark job that queries Salesforce

    Standing up an EMR cluster and writing a Spark connector job is heavy infrastructure and custom code for what is a simple SaaS extraction task — the opposite of the no-code, no-infrastructure requirement.

The concept

Amazon AppFlow is the managed, no-code ingestion service for SaaS sources (Salesforce, ServiceNow, Slack, and others), moving data to AWS targets like S3 and Redshift on a schedule or in response to events.

Why that’s the answer

AppFlow's built-in Salesforce connector plus scheduled flows delivers the nightly ingestion with pure configuration, satisfying both the no-custom-code and no-infrastructure constraints. Lambda and EMR both require custom integration code (and EMR adds cluster management), while DMS does not support SaaS APIs as sources.

How to reason it out
  1. Recognize the source as a SaaS application (Salesforce), which points to AppFlow rather than database-replication or generic ETL tools.
  2. Create an AppFlow connection to Salesforce with the appropriate credentials and scopes.
  3. Define a flow mapping Salesforce objects to an S3 destination, optionally with field mapping and filtering.
  4. Set the flow trigger to a nightly schedule.

Exam tip: For scheduled, no-code ingestion from SaaS applications like Salesforce, reach for Amazon AppFlow.

Data Ingestion on AWS: Kinesis, MSK, DMS, Glue, and Batch Patterns — the lesson that teaches this.

Question 5Data Ingestion and Transformation

A data engineer needs to run an existing AWS Glue ETL job every day at 02:00 UTC. There are no dependencies on other jobs and no complex workflow logic. Which scheduling approach has the LEAST operational overhead?

Choose one.

  • a
    Deploy Amazon MWAA and author an Airflow DAG that triggers the job daily

    MWAA runs a managed but always-on Airflow environment with meaningful hourly cost and operational surface; it is justified for complex multi-step DAGs with dependencies, not one standalone cron-style job.

  • b
    Create an Amazon EventBridge schedule with a cron expression that starts the Glue job Correct

    A serverless EventBridge cron schedule that targets the Glue job is pure configuration — no environment to run, patch, or pay for continuously — which is the right fit for a single independent job on a fixed time.

  • c
    Run a cron job on an Amazon EC2 instance that calls the StartJobRun API

    This introduces a server that must be patched, monitored, and kept highly available — a single point of failure and ongoing overhead a serverless scheduler avoids entirely.

  • d
    Create an AWS Step Functions state machine that loops with a Wait state until 02:00 UTC

    A perpetually looping state machine is an awkward, error-prone way to emulate cron; Step Functions itself is normally started by an EventBridge schedule rather than acting as one.

The concept

Time-based job scheduling on AWS is a spectrum: EventBridge cron schedules for simple time triggers, Glue workflows or Step Functions for dependency chains, and MWAA (Airflow) for complex, code-defined orchestration.

Why that’s the answer

With one job, no dependencies, and a fixed time, an EventBridge cron schedule is the minimal, serverless answer. MWAA carries continuous environment cost and administration for capability this task does not need, an EC2 cron host reintroduces server management, and a looping Step Functions machine misuses a workflow engine as a clock.

How to reason it out
  1. Assess workflow complexity: a single independent job needs a scheduler, not an orchestrator.
  2. Create an EventBridge schedule with the cron expression cron(0 2 * * ? *).
  3. Set the Glue StartJobRun action (or the Glue job target) as the schedule's target with an appropriate IAM role.
  4. Reserve MWAA or Step Functions for pipelines with branching, dependencies, or cross-service coordination.

Exam tip: For a simple time-based trigger of one job, use an EventBridge cron schedule — save Airflow and Step Functions for real orchestration needs.

Data Ingestion on AWS: Kinesis, MSK, DMS, Glue, and Batch Patterns — the lesson that teaches this.

Question 6Data Ingestion and Transformation

A data platform team migrated from an on-premises Hadoop stack and has hundreds of existing Apache Airflow DAGs orchestrating ingestion pipelines with complex cross-job dependencies, branching, and retries. They want to run these DAGs on AWS without rewriting them and without managing Airflow servers. Which service should they choose?

Choose one.

  • a
    Amazon EventBridge rules chained together for each DAG

    EventBridge is a scheduler and event router, not a DAG engine; recreating hundreds of Airflow DAGs with branching and retries as chained rules would be a full rewrite into a much less expressive tool.

  • b
    AWS Step Functions state machines translated from each DAG

    Step Functions is a capable orchestrator, but it uses its own state language — every DAG would have to be rewritten, violating the requirement to avoid rewriting existing Airflow code.

  • c
    Amazon Managed Workflows for Apache Airflow (MWAA) Correct

    MWAA runs open-source Apache Airflow as a managed service, so existing DAG code can be deployed largely as-is while AWS operates the scheduler, workers, and web server.

  • d
    Self-managed Airflow on an Amazon EC2 Auto Scaling group

    Running Airflow on EC2 preserves the DAGs but makes the team responsible for installing, patching, scaling, and upgrading Airflow — exactly the server management they want to avoid.

The concept

MWAA is managed open-source Airflow: it exists specifically so teams with an Airflow investment can keep their DAGs while offloading environment operations to AWS.

Why that’s the answer

The two constraints — reuse existing Airflow DAGs and avoid managing Airflow infrastructure — intersect only at MWAA. EventBridge and Step Functions would both force a rewrite into different orchestration models, and self-managed Airflow on EC2 keeps the DAGs but also keeps all the operational burden.

How to reason it out
  1. Identify the migration constraint: hundreds of existing Airflow DAGs must run without rewriting.
  2. Identify the operational constraint: no Airflow servers to manage.
  3. Choose MWAA, which runs Apache Airflow compatibly as a managed environment.
  4. Deploy DAGs by uploading them to the environment's S3 bucket, adjusting connections and operators for AWS services where needed.

Exam tip: When a team already lives in Airflow, MWAA delivers managed orchestration without a DAG rewrite.

Data Ingestion on AWS: Kinesis, MSK, DMS, Glue, and Batch Patterns — the lesson that teaches this.

What DEA-C01 domain 1 tests, topic by topic

The official exam guide breaks Data Ingestion and Transformation into 4 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published DEA-C01 practice questions per topic in Data Ingestion and Transformation
TopicWhat it coversQuestions
Perform data ingestionOfficial DEA-C01 task statement (guide v1.1). Reading streaming sources (Amazon Kinesis, Amazon MSK, DynamoDB Streams, AWS DMS, AWS Glue, Amazon Redshift) and batch sources (Amazon S3, Glue, Amazon EMR, DMS, Redshift, AWS Lambda, Amazon AppFlow); batch-ingestion configuration; consuming data APIs; setting up schedulers (Amazon EventBridge, Apache Airflow, time-based) for jobs and crawlers; event triggers (S3 Event Notifications, EventBridge); calling a Lambda function from Kinesis; IP allowlists; throttling and rate limits (DynamoDB, Amazon RDS, Kinesis); managing fan-in/fan-out for streaming distribution; replayability of ingestion pipelines; stateful vs stateless transactions.20
Transform and process dataOfficial DEA-C01 task statement. Optimizing container usage (Amazon EKS, ECS); connecting via JDBC/ODBC; integrating data from multiple sources; optimizing cost while processing; implementing transformation services by requirement (Amazon EMR, AWS Glue, Lambda, Amazon Redshift); converting formats (CSV to Apache Parquet); troubleshooting transformation failures and performance; creating data APIs; defining volume, velocity, and variety; integrating large language models (LLMs) for data processing.20
Orchestrate data pipelinesOfficial DEA-C01 task statement. Using orchestration services to build ETL workflows (Lambda, EventBridge, Amazon Managed Workflows for Apache Airflow [MWAA], AWS Step Functions, AWS Glue workflows); building pipelines for performance, availability, scalability, resiliency, and fault tolerance; implementing and maintaining serverless workflows; using notification services for alerts (Amazon SNS, Amazon SQS).20
Apply programming conceptsOfficial DEA-C01 task statement. Optimizing code to reduce runtime; configuring Lambda concurrency and performance; using languages and frameworks (Python, SQL, Scala, R, Java, Bash, PowerShell); software-engineering best practices (version control, testing, logging, monitoring); packaging and deploying serverless data pipelines with AWS SAM; using and mounting storage volumes from within Lambda functions; Infrastructure as Code (AWS CloudFormation, AWS CDK); CI/CD for data pipelines; distributed computing; data structures and algorithms.20
Total80

Revise Data Ingestion and Transformation before you drill it

Other DEA-C01 domains

Data Ingestion and Transformation: your questions

Data Ingestion and Transformation is domain 1 of the DEA-C01 exam guide and carries 34% of the scored content — the heaviest of the 4 domains. On a 65-question paper that works out to roughly 22 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official DEA-C01 exam guide.