SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
DVA-C02 · Domain 4

Troubleshooting and Optimization practice questions

Troubleshooting and Optimization is worth 18% of the DVA-C02 exam — the lightest of the 4 domains. Root-cause analysis, instrumenting code for observability, and optimizing applications on AWS. 6 fully worked examples are further down this page, answers included.

Exam weight
18%
the lightest of the 4 domains
Questions
60
across 3 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Troubleshooting and Optimization questions, fully explained

Questions from the DVA-C02 bank mapped to domain 4, with the answer key and the reasoning behind every option. None of them repeat the examples on the main DVA-C02 practice page.

Question 1Troubleshooting and Optimization

A developer suspects a Lambda function is timing out and wants the 20 most recent invocations whose logs contain "Task timed out", newest first. Which CloudWatch Logs Insights query achieves this?

Choose one.

  • a
    fields @timestamp, @message | filter @message like /Task timed out/ | sort @timestamp asc | limit 20

    sort asc orders oldest first, so with limit 20 this returns the 20 oldest matching events in the query window, not the most recent.

  • b
    fields @timestamp, @message | filter @message like /Task timed out/ | sort @timestamp desc | limit 20 Correct

    This selects the display fields, keeps only events matching the timeout text, orders newest first with sort desc, and caps the output at 20 — exactly the stated goal.

  • c
    stats count() by bin(5m) | filter @message like /Task timed out/

    This aggregates event counts into time buckets rather than returning the individual events, and filtering must narrow the events before aggregation, not after it.

  • d
    parse @message /Task timed out/ | limit 20

    parse extracts fields from message text; it does not restrict results to matching events, so this returns the first 20 events of any kind.

The concept

Logs Insights queries chain commands with pipes: filter narrows events, sort orders them, and limit caps how many return. The system fields @timestamp and @message are always available, and sort @timestamp desc is how "most recent first" is expressed.

Why that’s the answer

Option b is the canonical shape for "the latest N events matching a pattern": filter to the timeout text, sort descending by timestamp, limit to 20. Option a sorts ascending and therefore returns the oldest matches. Option c produces bucketed counts, not events, and applies its filter in the wrong position. Option d uses parse, which extracts fields rather than filtering, so non-matching events still flow through.

How to reason it out
  1. Choose the function's log group and a time range covering the incident.
  2. filter @message like /Task timed out/ to keep only the relevant events.
  3. sort @timestamp desc so the newest occurrences come first.
  4. limit 20 to cap the result set at the number you need.

Exam tip: filter + sort desc + limit is the Logs Insights pattern for "the most recent N matching events".

Root Cause Analysis on AWS: Reading Logs, Metrics, and Error Codes — the lesson that teaches this.

Question 2Troubleshooting and Optimization

A REST API using API Gateway with a Lambda proxy integration intermittently returns 502 Bad Gateway. The Lambda function's Errors metric is flat at zero and its Duration is normal. What is the most likely cause?

Choose one.

  • a
    The function is running past API Gateway's 29-second integration timeout

    An integration that exceeds the timeout produces a 504 Gateway Timeout, not a 502, and Duration is reported as normal here.

  • b
    One code path returns a response that is missing statusCode or has a body that is not a string Correct

    With proxy integration the function must return a numeric statusCode and a string body; a raw object or missing statusCode cannot be mapped to an HTTP response, so API Gateway emits 502 even though the function succeeded.

  • c
    Clients have exhausted their usage plan quota

    Quota exhaustion and throttling return 429 Too Many Requests, never 502.

  • d
    An AWS WAF rule is blocking a subset of the requests

    WAF blocks surface as 403 Forbidden; the request never reaches the integration, so it cannot produce a bad-gateway error.

The concept

With Lambda proxy integration, API Gateway performs no transformation: the function must return an envelope with a numeric statusCode, optional headers, a body that is a string (typically JSON.stringify output), and isBase64Encoded for binary. Any response that breaks this contract produces a 502.

Why that’s the answer

The tell is 502 with a healthy function: the function completed (Errors is zero, Duration normal), yet the gateway could not translate its response — the signature of a malformed proxy envelope on some code path. A timeout would show as 504 and abnormal duration, quota exhaustion returns 429, and a WAF block returns 403 before the integration ever runs.

How to reason it out
  1. Confirm from metrics that the function itself reports no errors and normal duration.
  2. Query the function's logs for the failing request IDs and inspect what each code path returned.
  3. Find the branch returning a raw object, omitting statusCode, or passing an unserialized body.
  4. Fix that path to return the full envelope with a stringified body, and add a test for it.

Exam tip: 502 from API Gateway with a healthy Lambda function almost always means a malformed proxy response.

Root Cause Analysis on AWS: Reading Logs, Metrics, and Error Codes — the lesson that teaches this.

Question 3Troubleshooting and Optimization

A Lambda function behind an API Gateway REST API has its function timeout set to 60 seconds. Requests that need about 45 seconds of processing fail with 504 Gateway Timeout, yet the function's logs show those invocations completing successfully. What explains the failures?

Choose one.

  • a
    Lambda functions cannot be configured with timeouts longer than 30 seconds

    Lambda supports timeouts up to 15 minutes; the 60-second function timeout is valid — it is the gateway in front that stops waiting.

  • b
    The integration exceeded API Gateway's default 29-second integration timeout, so the gateway gave up before the function finished Correct

    API Gateway's default integration timeout is 29 seconds regardless of the function's own timeout; the gateway returns 504 while the function keeps running to completion in the background.

  • c
    The function is returning a malformed proxy response envelope

    A malformed proxy response causes a 502 Bad Gateway, not a 504, and it would not correlate with request duration.

  • d
    The API's usage plan is throttling long-running requests

    Throttling returns 429 Too Many Requests and is based on request rate, not on how long an individual request takes.

The concept

A 504 Gateway Timeout from API Gateway means the integration did not respond within the gateway's integration timeout — 29 seconds by default — which is independent of, and can be far shorter than, the backend's own timeout setting.

Why that’s the answer

The function completes in its logs because its own 60-second timeout allows the work, but API Gateway stopped waiting at the 29-second default and returned 504 to the client. Option a misstates Lambda limits (up to 15 minutes). Option c describes the 502 failure mode. Option d confuses duration with rate — throttling is 429 and unrelated to how long one request runs.

How to reason it out
  1. Match the 504 responses against the function's logs and confirm those invocations ran longer than 29 seconds.
  2. Recognize that the gateway's integration timeout, not the function timeout, is the binding limit.
  3. Shorten the synchronous work, or move long processing to an asynchronous pattern that responds immediately and completes in the background.
  4. Reserve raising timeouts as a last resort — a client waiting tens of seconds is usually a design smell.

Exam tip: 504 means the integration blew through API Gateway's 29-second default timeout — the function's own timeout is irrelevant to the client.

Root Cause Analysis on AWS: Reading Logs, Metrics, and Error Codes — the lesson that teaches this.

Question 4Troubleshooting and Optimization

Users of a mobile app report that requests to an API Gateway API start failing with HTTP 429 late each afternoon and succeed again the next morning. What is the most likely cause?

Choose one.

  • a
    A Lambda authorizer is rejecting the requests

    An authorizer denial returns 403 Forbidden, not 429, and would not follow a time-of-day reset pattern.

  • b
    The requests are failing API Gateway request validation

    Failed request validation returns 400 Bad Request and depends on payload shape, not on how much traffic has already been sent that day.

  • c
    The API keys' usage plan quota is being exhausted partway through each day Correct

    429 Too Many Requests signals throttling or quota exhaustion; a daily quota that runs out in the afternoon and resets overnight matches the recurring pattern exactly.

  • d
    The backend integration is timing out under afternoon load

    An integration timeout surfaces as 504 Gateway Timeout; it would not produce 429 responses.

The concept

HTTP 429 Too Many Requests from API Gateway always means throttling: a rate or burst limit was hit, or a usage plan quota was exhausted. It is never an authentication or validation failure.

Why that’s the answer

The daily rhythm — failures beginning each afternoon and clearing overnight — is the signature of a per-day usage plan quota being consumed and then reset. An authorizer problem would return 403 and show no daily cycle, validation failures return 400 and depend on the payload, and a timing-out backend returns 504.

How to reason it out
  1. Read the status code first: 429 narrows the cause to throttling or quota before any log diving.
  2. Check the usage plan attached to the affected API keys and compare its daily quota against actual request volume.
  3. Confirm in the stage's metrics when throttled requests begin each day.
  4. Raise the quota, spread traffic, or move heavy clients to a plan that matches their volume.

Exam tip: 429 is always throttling or quota — a failure pattern that resets daily points at a usage plan quota.

Root Cause Analysis on AWS: Reading Logs, Metrics, and Error Codes — the lesson that teaches this.

Question 5Troubleshooting and Optimization

After a security review, a subset of clients suddenly receives HTTP 403 Forbidden from an API Gateway API. The stage's throttling metrics show no throttle events, and the backend Lambda function's logs show no invocations for the failing requests. Where should the developer look for the cause?

Choose one.

  • a
    The clients' request rate against the stage's rate and burst limits

    Exceeding rate or burst limits returns 429, and the throttling metrics already show no throttle events.

  • b
    The Lambda function's response envelope for a malformed statusCode or body

    A malformed proxy response causes 502, and it cannot be the cause here because the function is never invoked for the failing requests.

  • c
    The authorization layer: an AWS WAF web ACL rule, a resource policy, an authorizer, or a missing or invalid API key Correct

    403 is always an authorization-shaped rejection, and the request being blocked before the integration runs explains why the function logs nothing.

  • d
    The integration timeout configuration on the affected routes

    A timed-out integration returns 504; it also requires the request to reach the backend, which these requests never do.

The concept

403 Forbidden from API Gateway is always an authorization problem — an IAM or Lambda authorizer denial, a resource policy, a missing or invalid API key, or an AWS WAF block. It is never throttling, which returns 429.

Why that’s the answer

Two clues triangulate the answer: the 403 code itself points at auth, and the absence of backend invocations proves the request was rejected before the integration — where authorizers, resource policies, API key checks, and WAF sit. The timing after a security review makes a new WAF rule or tightened policy the natural suspect. Throttling is ruled out by both the code and the metrics, and 502/504 causes require the backend to be involved.

How to reason it out
  1. Confirm the status code is 403 and note that throttling metrics are clean.
  2. List what changed recently — here, a security review likely added WAF rules or tightened policies.
  3. Check the WAF web ACL's blocked-request metrics and sampled requests for the affected clients.
  4. Review the API's resource policy, authorizer behavior, and the clients' API keys until the denying control is found.

Exam tip: 403 means an authorization control said no — check WAF, policies, authorizers, and API keys, never traffic volume.

Root Cause Analysis on AWS: Reading Logs, Metrics, and Error Codes — the lesson that teaches this.

Question 6Troubleshooting and Optimization

An API Gateway REST API validates request bodies against a JSON schema model. Immediately after a new version of the client app ships, a large share of requests fails with HTTP 400 and the backend shows no corresponding invocations. What is the most likely cause?

Choose one.

  • a
    The new app version sends payloads that fail the API's request validation Correct

    400 Bad Request with no backend invocation is exactly how enabled request validation rejects malformed payloads at the gateway, and the correlation with the app release points at a changed request shape.

  • b
    The backend is returning malformed responses for the new payloads

    A malformed backend response produces 502, and the backend is never invoked for these requests, so it cannot be the source.

  • c
    The new app version's traffic is exceeding the stage's burst limit

    Burst and rate limit violations return 429 Too Many Requests, not 400.

  • d
    The app is calling a stage whose cache is serving stale responses

    A stage cache serves previously cached successful responses; it does not generate 400 errors for incoming requests.

The concept

HTTP 400 Bad Request means the client sent something the API refuses to accept — with API Gateway request validation enabled, payloads that fail the JSON schema model are rejected at the gateway before any integration runs.

Why that’s the answer

The combination of 400s, no backend invocations, and onset at the moment a new client shipped identifies validation rejecting the new payload shape. A backend response problem (b) would be a 502 and requires invocations that never happen; throttling (c) is 429; and a stage cache (d) replays cached successes rather than minting client errors.

How to reason it out
  1. Correlate the 400 spike's start time with the client release.
  2. Check the gateway's execution logs for the validation error messages on rejected requests.
  3. Diff the new app's request payloads against the JSON schema model.
  4. Fix the client payload, or version the model if the API contract legitimately changed.

Exam tip: 400 with no backend invocation means the request died at validation — diff the payload against the schema.

Root Cause Analysis on AWS: Reading Logs, Metrics, and Error Codes — the lesson that teaches this.

What DVA-C02 domain 4 tests, topic by topic

The official exam guide breaks Troubleshooting and Optimization into 3 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published DVA-C02 practice questions per topic in Troubleshooting and Optimization
TopicWhat it coversQuestions
Assist in a root cause analysisExam guide task 4.1. Debugging code and service-integration issues; interpreting and querying metrics, logs, and traces; custom metrics (e.g. CloudWatch embedded metric format); dashboards and insights for application health; troubleshooting deployment failures from service logs.20
Instrument code for observabilityExam guide task 4.2. Logging vs monitoring vs observability; effective and structured logging strategies; emitting custom metrics from code; distributed tracing with annotations (AWS X-Ray); notification alerts (quota limits, deployment completions); health checks and readiness probes.20
Optimize applications by using AWS services and featuresExam guide task 4.3. Concurrency and profiling; determining minimum memory/compute sizing; subscription filter policies for messaging; content caching by request headers and application-level caching; optimizing resource usage, analyzing performance issues, and finding bottlenecks through application logs.20
Total60

Revise Troubleshooting and Optimization before you drill it

Other DVA-C02 domains

Troubleshooting and Optimization: your questions

Troubleshooting and Optimization is domain 4 of the DVA-C02 exam guide and carries 18% of the scored content — the lightest of the 4 domains. On a 65-question paper that works out to roughly 12 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official DVA-C02 exam guide.