SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
DEA-C01 · Domain 3

Data Operations and Support practice questions

Data Operations and Support is worth 22% of the DEA-C01 exam — the 3rd-heaviest of the 4 domains. Automating data processing, analyzing data, maintaining and monitoring pipelines, and ensuring data quality. Official weighting 22%. 6 fully worked examples are further down this page, answers included.

Exam weight
22%
the 3rd-heaviest of the 4 domains
Questions
80
across 4 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Data Operations and Support questions, fully explained

Questions from the DEA-C01 bank mapped to domain 3, with the answer key and the reasoning behind every option. None of them repeat the examples on the main DEA-C01 practice page.

Question 1Data Operations and Support

A data engineer needs to invoke a Lambda function exactly once at 6:00 PM local time on the last day of the current quarter to trigger a one-time data export, and the schedule must respect the Europe/London time zone including daylight saving changes. Which approach requires the least operational effort?

Choose one.

  • a
    Create an Amazon EventBridge rule with a cron expression

    EventBridge rules only evaluate cron expressions in UTC, have no one-time schedule concept, and would keep firing every quarter unless someone remembers to delete the rule.

  • b
    Create a one-time schedule in Amazon EventBridge Scheduler with the Europe/London time zone Correct

    Correct. EventBridge Scheduler supports one-time schedules with an explicit time zone, handles daylight saving automatically, and can invoke Lambda directly.

  • c
    Run an EC2 instance with an at job configured for the target time

    Standing up and patching an EC2 instance to run a single scheduled command is the highest-effort option and introduces an unnecessary single point of failure.

  • d
    Use a Step Functions Wait state that pauses until the target timestamp

    A Step Functions execution parked in a long Wait state works but is a workaround: you pay for and monitor a running execution for weeks, and time zone or DST logic must be computed before starting it.

The concept

Amazon EventBridge Scheduler is the purpose-built scheduling service, and one-time schedules with time-zone awareness are exactly the capabilities that distinguish it from EventBridge rules.

Why that’s the answer

EventBridge Scheduler supports at() expressions for one-time invocations, lets you set an IANA time zone such as Europe/London so daylight saving transitions are handled for you, and integrates directly with Lambda as a target. EventBridge scheduled rules cannot do one-time schedules and evaluate cron only in UTC, so DST would silently shift the invocation time. The EC2 and Step Functions options both work mechanically but carry ongoing cost and operational burden for what should be a fire-and-forget schedule.

How to reason it out
  1. Note the two constraints: a one-time invocation and local time zone plus DST correctness.
  2. Recall that EventBridge Scheduler supports at() one-time expressions and IANA time zones, while EventBridge rules support neither.
  3. Create the schedule with the Lambda function as the target and let Scheduler handle delivery and retries.

Exam tip: One-time schedules and time-zone-aware schedules are EventBridge Scheduler features; classic EventBridge rules are recurring and UTC-only.

Automate Data Processing on AWS: Step Functions, MWAA, Lambda & EventBridge — the lesson that teaches this.

Question 2Data Operations and Support

A data engineer uploads a new DAG file to the S3 bucket of an Amazon MWAA environment, but after 30 minutes the DAG does not appear in the Airflow UI. Other DAGs in the environment run normally. What should the engineer check first to find the cause?

Choose one.

  • a
    The Amazon MWAA worker logs for failed task executions

    Worker logs capture task execution output. A DAG that is not visible in the UI has never been scheduled or executed, so worker logs will contain nothing about it.

  • b
    The AWS CloudTrail event history for the S3 PutObject call

    CloudTrail confirms the upload API call occurred, but the file being in S3 is not in doubt; the problem is that Airflow cannot parse or import it, which CloudTrail cannot show.

  • c
    The Airflow scheduler logs and DAG processing logs in Amazon CloudWatch for import or parse errors Correct

    Correct. MWAA publishes scheduler and DAG processing logs to CloudWatch; a DAG that never appears in the UI almost always failed to parse (syntax error, missing import, or bad top-level code), and the parse error is recorded there.

  • d
    The Amazon MWAA web server access logs

    Web server logs record UI and API access to Airflow. They say nothing about whether the scheduler successfully parsed a new DAG file.

The concept

In MWAA, the scheduler continuously parses DAG files from S3; a DAG missing from the UI is a parse-time failure, and MWAA surfaces parse errors through scheduler and DAG processing logs in CloudWatch.

Why that’s the answer

MWAA syncs the dags/ prefix of the environment bucket and the Airflow scheduler attempts to import each Python file. If the file has a syntax error, references a package that is not in requirements.txt, or raises at import time, the DAG never registers and therefore never appears in the UI. Those failures are written to the scheduler log group (and DAG processing logs) in CloudWatch, which is why enabling and reading those logs is the standard first troubleshooting step. Worker logs only exist for tasks that actually ran, CloudTrail only proves the object was uploaded, and web server logs cover UI traffic.

How to reason it out
  1. Confirm the DAG file landed under the correct dags/ prefix in the environment S3 bucket.
  2. Open the CloudWatch scheduler and DAG processing log groups for the MWAA environment and search for the DAG file name.
  3. Fix the reported import or syntax error (often a dependency missing from requirements.txt) and wait for the scheduler to re-parse.

Exam tip: A DAG that never shows up in the MWAA UI failed to parse; look in the CloudWatch scheduler and DAG processing logs.

Automate Data Processing on AWS: Step Functions, MWAA, Lambda & EventBridge — the lesson that teaches this.

Question 3Data Operations and Support

A Step Functions Standard workflow that orchestrates a data pipeline failed overnight. The on-call data engineer needs to determine exactly which state failed, what input that state received, and the error message it returned. Where should the engineer look?

Choose one.

  • a
    The AWS CloudTrail data events for the state machine

    CloudTrail records the management API calls such as StartExecution, not the internal state-by-state transitions, inputs, or error causes of an execution.

  • b
    The execution history of the failed execution in the Step Functions console Correct

    Correct. Standard workflow executions record a complete event-by-event history: each state transition, its input and output, and the error and cause for the failed state, viewable in the console or via GetExecutionHistory.

  • c
    Amazon CloudWatch metrics for the state machine

    CloudWatch metrics show aggregate counts such as ExecutionsFailed. They reveal that a failure happened, not which state failed or with what input and error.

  • d
    The Amazon EventBridge event bus default archive

    EventBridge archives only capture events sent to a bus with an archive configured; Step Functions state transitions are not published there by default and would lack input and error detail anyway.

The concept

Step Functions Standard workflows persist a full execution history that is the primary troubleshooting tool for failed pipelines.

Why that’s the answer

Every Standard execution logs an ordered event stream: state entered, state exited, task scheduled, task failed, and so on, with the JSON input and output of each state and the error name plus cause for failures. The console visualizes this with the failed state highlighted, and the GetExecutionHistory API returns the same data programmatically. That directly answers all three questions in the stem: which state, what input, what error. CloudWatch metrics are aggregates, CloudTrail covers control-plane API calls, and EventBridge is not a store of state-transition detail.

How to reason it out
  1. Open the state machine in the Step Functions console and select the failed execution.
  2. Use the graph view to spot the state highlighted in red, then open its events to read the input, output, error, and cause fields.
  3. Optionally call GetExecutionHistory from the CLI or SDK to pull the same event detail for automation or ticketing.

Exam tip: Debug failed Step Functions executions with the per-execution history, which records every state transition with inputs, outputs, and error causes.

Automate Data Processing on AWS: Step Functions, MWAA, Lambda & EventBridge — the lesson that teaches this.

Question 4Data Operations and Support

A Lambda function must run SQL statements against an Amazon Redshift provisioned cluster after each ETL load. The team wants to avoid managing JDBC drivers, persistent database connections, and VPC networking between Lambda and the cluster. Which approach meets these requirements?

Choose one.

  • a
    Open a JDBC connection from Lambda to the cluster endpoint

    JDBC is exactly what the team wants to avoid: it requires driver management, connection handling, and Lambda VPC configuration to reach the cluster endpoint.

  • b
    Call the Amazon Redshift Data API from the Lambda function Correct

    Correct. The Redshift Data API is an HTTPS-based API (boto3 redshift-data client) that runs SQL asynchronously without drivers, connection pools, or the Lambda function being in the cluster VPC.

  • c
    Use Amazon Redshift Spectrum to run the statements

    Redshift Spectrum is a feature for querying external data in S3 from within Redshift SQL; it is not a connectivity mechanism for external applications to submit SQL.

  • d
    Configure an Amazon AppFlow flow to execute the SQL

    AppFlow moves data between SaaS applications and AWS services; it does not execute arbitrary SQL statements against a Redshift cluster on demand.

The concept

The Amazon Redshift Data API lets applications run SQL over HTTPS with IAM or Secrets Manager authentication, removing the need for drivers and persistent connections.

Why that’s the answer

With the Data API, the Lambda function calls ExecuteStatement on the redshift-data client, receives a statement ID immediately, and later fetches results with GetStatementResult or reacts to the completion event. Because the API is a regional HTTPS endpoint, the function does not need to be attached to the cluster VPC, and there are no JDBC/ODBC drivers or connection pools to manage, which is a major fit for short-lived, event-driven callers like Lambda. The asynchronous model also sidesteps Lambda timeout pressure for long statements.

How to reason it out
  1. Grant the Lambda execution role redshift-data permissions and access to the cluster credentials via Secrets Manager or temporary IAM credentials.
  2. Call ExecuteStatement (or BatchExecuteStatement) with the SQL, cluster identifier, and database.
  3. Poll DescribeStatement or subscribe to the Data API completion event, then read results with GetStatementResult if needed.

Exam tip: For serverless or driverless SQL access to Redshift, especially from Lambda, use the Redshift Data API.

Automate Data Processing on AWS: Step Functions, MWAA, Lambda & EventBridge — the lesson that teaches this.

Question 5Data Operations and Support

A Python Lambda function must start an AWS Glue ETL job named daily-sales-etl and pass the run date as a job argument. Which boto3 call accomplishes this?

Choose one.

  • a
    glue_client.create_job(Name='daily-sales-etl', Arguments={'--run_date': '2026-07-29'})

    create_job defines a new job (script location, role, capacity); it does not execute an existing job, and calling it for an existing name fails.

  • b
    glue_client.start_crawler(Name='daily-sales-etl')

    start_crawler runs a Glue crawler that catalogs data; it cannot run an ETL job or accept job arguments.

  • c
    glue_client.start_job_run(JobName='daily-sales-etl', Arguments={'--run_date': '2026-07-29'}) Correct

    Correct. The Glue client exposes start_job_run, which takes the job name and an Arguments map of job parameters (keys prefixed with --) and returns a JobRunId.

  • d
    lambda_client.invoke(FunctionName='daily-sales-etl')

    lambda_client.invoke calls another Lambda function. A Glue job is not a Lambda function and cannot be started this way.

The concept

Automating AWS services from code means calling the right SDK operation; for Glue ETL jobs the run operation is start_job_run with an Arguments dictionary.

Why that’s the answer

boto3 mirrors the Glue API: StartJobRun starts an existing job definition and returns a JobRunId that can be tracked with get_job_run. Runtime parameters are passed in the Arguments map, and Glue convention prefixes each key with two dashes so the script can read it via getResolvedOptions. create_job is a definition-time call, start_crawler targets crawlers rather than jobs, and lambda_client.invoke addresses Lambda functions only.

How to reason it out
  1. Create a Glue client with boto3.client('glue') inside the Lambda function.
  2. Call start_job_run with JobName and an Arguments map whose keys use the -- prefix.
  3. Capture the returned JobRunId and monitor the run with get_job_run or a downstream EventBridge job state-change event.

Exam tip: start_job_run (with an Arguments map of --prefixed parameters) is the boto3 call that executes an existing Glue job.

Automate Data Processing on AWS: Step Functions, MWAA, Lambda & EventBridge — the lesson that teaches this.

Question 6Data Operations and Support

A pipeline consists entirely of AWS Glue components: a crawler that updates the Data Catalog, followed by three Glue ETL jobs that must run in sequence, with the whole chain triggered on a schedule. The team wants to stay within Glue and monitor the chain as a single unit. Which feature should they use?

Choose one.

  • a
    An AWS Glue workflow with triggers connecting the crawler and jobs Correct

    Correct. Glue workflows natively chain crawlers and jobs with scheduled and conditional triggers and present the whole pipeline as one monitorable graph inside Glue.

  • b
    An Amazon MWAA environment with a new DAG

    MWAA can orchestrate Glue, but provisioning and paying for a managed Airflow environment for an all-Glue linear chain contradicts the requirement to stay within Glue.

  • c
    Four separate EventBridge Scheduler schedules with staggered start times

    Staggered schedules guess at durations rather than reacting to completion; a slow crawler would overlap with the first job, and there is no single unit to monitor.

  • d
    An AWS Glue DataBrew project

    DataBrew is a visual data preparation tool for building recipes; it does not orchestrate crawlers and ETL jobs into a dependency chain.

The concept

AWS Glue workflows are the built-in orchestration layer for Glue-native pipelines, wiring crawlers and jobs together with triggers and exposing a single run view.

Why that’s the answer

A Glue workflow starts from a scheduled trigger, runs the crawler, and uses conditional (on-success) triggers to fire each ETL job only after the previous component succeeds. Workflow run properties can even pass state between jobs. Because the entire chain lives in Glue, the console shows one graph with per-node status, satisfying the single-unit monitoring requirement without introducing another orchestration service. MWAA and Step Functions are appropriate when the pipeline spans services beyond Glue; staggered timers are brittle; DataBrew solves a different problem.

How to reason it out
  1. Create a Glue workflow and add a scheduled trigger as its starting point.
  2. Add the crawler, then chain the three jobs with on-success conditional triggers.
  3. Monitor workflow runs from the Glue console graph and alert on failed nodes.

Exam tip: For pipelines made only of Glue crawlers and jobs, Glue workflows with triggers provide orchestration without an external orchestrator.

Automate Data Processing on AWS: Step Functions, MWAA, Lambda & EventBridge — the lesson that teaches this.

What DEA-C01 domain 3 tests, topic by topic

The official exam guide breaks Data Operations and Support into 4 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published DEA-C01 practice questions per topic in Data Operations and Support
TopicWhat it coversQuestions
Automate data processing by using AWS servicesOfficial DEA-C01 task statement. Orchestrating pipelines (Amazon MWAA, AWS Step Functions); troubleshooting managed workflows; calling SDKs; using service features to process data (Amazon EMR, Redshift, AWS Glue); consuming and maintaining data APIs; preparing data for transformation (AWS Glue DataBrew, Amazon SageMaker Unified Studio); querying data (Amazon Athena); using AWS Lambda to automate processing; managing events and schedulers (Amazon EventBridge).20
Analyze data by using AWS servicesOfficial DEA-C01 task statement. Visualizing data (AWS Glue DataBrew, Amazon QuickSight); verifying and cleaning data (Lambda, Athena, QuickSight, Jupyter notebooks, Amazon SageMaker Data Wrangler); using SQL in Amazon Redshift and Athena to query and create views; using Athena notebooks with Apache Spark; tradeoffs between provisioned and serverless services; data aggregation, rolling averages, grouping, and pivoting.20
Maintain and monitor data pipelinesOfficial DEA-C01 task statement. Extracting logs for audits; deploying logging and monitoring for traceability; using notifications to send alerts; troubleshooting performance; tracking API calls with AWS CloudTrail; troubleshooting and maintaining pipelines (AWS Glue, Amazon EMR); using Amazon CloudWatch Logs; analyzing logs (Athena, EMR, Amazon OpenSearch Service, CloudWatch Logs Insights).20
Ensure data qualityOfficial DEA-C01 task statement. Running data quality checks while processing (for example, checking for empty fields); defining data quality rules (AWS Glue DataBrew); investigating data consistency; data sampling techniques; implementing data skew mechanisms.20
Total80

Revise Data Operations and Support before you drill it

Other DEA-C01 domains

Data Operations and Support: your questions

Data Operations and Support is domain 3 of the DEA-C01 exam guide and carries 22% of the scored content — the 3rd-heaviest of the 4 domains. On a 65-question paper that works out to roughly 14 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official DEA-C01 exam guide.