SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
SAA-C03 · Domain 2

Design Resilient Architectures practice questions

Design Resilient Architectures is worth 26% of the SAA-C03 exam — the 2nd-heaviest of the 4 domains. Scalable, loosely coupled designs and highly available, fault-tolerant architectures measured against RTO and RPO. 6 fully worked examples are further down this page, answers included.

Exam weight
26%
the 2nd-heaviest of the 4 domains
Questions
40
across 2 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Design Resilient Architectures questions, fully explained

Questions from the SAA-C03 bank mapped to domain 2, with the answer key and the reasoning behind every option. None of them repeat the examples on the main SAA-C03 practice page.

Question 1Design Resilient Architectures

An order-intake API calls a downstream fulfillment service synchronously for every order. During flash promotions the fulfillment service is overwhelmed, calls fail, and orders are permanently lost. The company requires that no accepted order is ever lost and that fulfillment capacity grows automatically with demand. Which combination of steps should a solutions architect take? (Select TWO.)

Choose TWO.

  • a
    Change the API to send each accepted order to an Amazon SQS queue instead of calling the fulfillment service directly. Correct

    The queue durably persists every accepted order until fulfillment processes and deletes it, so a saturated or failed fulfillment service can no longer cause loss.

  • b
    Publish orders to an Amazon SNS topic and subscribe the fulfillment instances as HTTP endpoints.

    SNS push delivery retries for a while and then drops undeliverable messages; overwhelmed endpoints can still lose orders because there is no consumer-paced durable buffer.

  • c
    Run the fulfillment service in an Auto Scaling group that scales based on the depth of the order queue. Correct

    Scaling on queue depth grows fulfillment capacity exactly when a backlog forms and shrinks it afterward, meeting the automatic-growth requirement.

  • d
    Configure the API to retry failed fulfillment calls with exponential backoff.

    Retries help transient errors, but the order exists only in the API instance memory while retrying; sustained saturation or an API crash mid-retry still loses it.

  • e
    Migrate the fulfillment database to a larger instance class.

    This guesses at a bottleneck the stem never identifies and does nothing to decouple the tiers or prevent loss when the service itself saturates.

The concept

Guaranteeing no lost work between a producer and an overwhelmed consumer requires two things together: a durable buffer that owns each message until it is processed, and a consumer tier that scales on the buffer's depth.

Why that’s the answer

The queue and the queue-driven Auto Scaling group are two halves of one design. SQS satisfies the no-loss requirement because a message survives consumer failure and reappears after the visibility timeout until a worker deletes it post-success. Scaling fulfillment on queue depth satisfies the elastic-capacity requirement without human action. SNS fails the no-loss requirement on its own: it is push-based, retains nothing for consumer-paced retrieval, and exhausts retries against a saturated endpoint. Client-side exponential backoff keeps the tight coupling; the order lives in volatile memory during retries, so the failure window is narrowed but not closed, and it adds latency to the intake path during exactly the traffic peak. Enlarging the database is vertical scaling aimed at an unstated bottleneck; the stem describes a coupling failure, not a database failure.

How to reason it out
  1. Restate the two hard requirements: zero lost orders and automatic capacity growth.
  2. Map zero loss to durable buffering: only SQS holds an order independently of both the API and fulfillment being healthy.
  3. Map automatic growth to a scaling signal: queue depth is the direct measure of unprocessed fulfillment work.
  4. Eliminate SNS push and client retries because both leave a window where the only copy of an order can vanish.
  5. Eliminate database resizing because it neither buffers nor scales the failing tier.

Exam tip: No-lost-work plus elastic capacity is always the pair: SQS as the durable buffer, workers scaling on queue depth.

Scalable, Loosely Coupled Architectures: SQS, SNS, EventBridge, Step Functions — the lesson that teaches this.

Question 2Design Resilient Architectures

A banking platform queues payment instructions for asynchronous processing. Instructions for any given account must be applied in the exact order they were submitted and must never be processed twice, while instructions for different accounts may be processed in parallel. Which messaging design meets these requirements?

Choose one.

  • a
    An SQS standard queue with consumers written to be idempotent.

    Idempotency neutralizes duplicates, but a standard queue offers only best-effort ordering, so per-account sequence cannot be guaranteed.

  • b
    An SQS standard queue with long polling enabled.

    Long polling reduces empty receives and cost; it has no effect on ordering or duplicate delivery.

  • c
    An SNS standard topic with the processing service as a subscriber.

    Standard SNS does not guarantee ordering, delivery retries can produce duplicates, and there is no durable consumer-paced buffer for backlog processing.

  • d
    An SQS FIFO queue with the account ID used as the message group ID. Correct

    FIFO guarantees strict ordering within each message group and exactly-once processing within the deduplication window, and per-account group IDs let different accounts process in parallel.

The concept

SQS FIFO queues exist for exactly two guarantees standard queues cannot make: strict ordering within a message group and exactly-once processing within the deduplication window. The message group ID scopes ordering, enabling parallelism across groups.

Why that’s the answer

The scenario states both FIFO trigger conditions explicitly: strict per-account order and no duplicates. Using the account ID as the message group ID gives each account its own ordered lane while different accounts process concurrently, which also addresses the parallelism requirement. The idempotent-consumer option is the strongest distractor because idempotency genuinely solves the duplicate half, but it cannot restore ordering that a standard queue never promised, so it fails half the requirement. Long polling is a cost and latency optimization for retrieving messages; it is orthogonal to delivery semantics. Standard SNS fails on all counts: unordered, at-least-once push delivery with no queueing for the consumer to pace itself.

How to reason it out
  1. Extract the requirements: strict order per account, no duplicate processing, parallelism across accounts.
  2. Match strict ordering and exactly-once to FIFO, the only queue type that provides them.
  3. Scope ordering with the message group ID set to the account ID so unrelated accounts do not serialize behind each other.
  4. Reject standard-queue options because best-effort ordering violates the per-account sequence requirement no matter how consumers are written.

Exam tip: Order plus no-duplicates means FIFO, and the message group ID is how you get per-entity ordering with cross-entity parallelism.

Scalable, Loosely Coupled Architectures: SQS, SNS, EventBridge, Step Functions — the lesson that teaches this.

Question 3Design Resilient Architectures

An IoT platform ingests several hundred thousand telemetry messages per minute from millions of sensors. Each message is independent, arrival order does not matter, and the consumers are already idempotent. A team member proposes a FIFO queue to be safe. Which queue design meets the throughput requirement with the LEAST complexity?

Choose one.

  • a
    A single SQS FIFO queue.

    FIFO trades throughput for ordering and deduplication guarantees this workload does not need, introducing a capped-throughput bottleneck for zero benefit.

  • b
    Multiple SQS FIFO queues with a routing function that shards messages across them.

    Sharding around the FIFO cap can reach the throughput, but it adds a routing layer and multiple queues to operate, all to preserve guarantees the scenario explicitly does not require.

  • c
    An Amazon MQ broker cluster.

    Amazon MQ targets migrations of existing broker-protocol applications; it requires broker sizing and management and does not match SQS standard's effectively unlimited scale for a new build.

  • d
    A single SQS standard queue. Correct

    Standard queues offer nearly unlimited throughput with at-least-once delivery, and idempotent consumers already neutralize the occasional duplicate, so nothing more is needed.

The concept

Default to SQS standard and choose FIFO only when the workload genuinely requires strict ordering or exactly-once processing. FIFO buys those guarantees by capping throughput; standard queues scale nearly without limit.

Why that’s the answer

The scenario removes both reasons FIFO exists: order does not matter, and duplicates are already handled by idempotent consumers. That leaves throughput and simplicity as the deciding factors, and a single standard queue delivers effectively unlimited throughput with nothing extra to build or operate. The single FIFO queue is the safety-blanket trap the stem sets up: choosing guarantees you do not need buys a hard throughput ceiling. The sharded-FIFO option is the technically-workable-but-inferior design; it reintroduces the throughput via complexity, adding a routing layer, more queues, and more failure modes, which loses under LEAST complexity. Amazon MQ is for lift-and-shift of applications speaking broker protocols like AMQP or JMS; for a cloud-native firehose it brings broker capacity management without matching standard SQS scale.

How to reason it out
  1. Check the FIFO triggers: no ordering requirement and duplicates already tolerated, so FIFO is unjustified.
  2. Recall the cost of FIFO: capped throughput versus nearly unlimited for standard.
  3. Compare the workable options against the qualifier: one standard queue versus a sharded FIFO fleet with routing code; the simpler design wins.
  4. Confirm consumers being idempotent makes at-least-once delivery fully acceptable.

Exam tip: Picking FIFO without an ordering or dedup requirement buys a bottleneck; standard plus idempotent consumers is the high-throughput default.

Scalable, Loosely Coupled Architectures: SQS, SNS, EventBridge, Step Functions — the lesson that teaches this.

Question 4Design Resilient Architectures

Workers pull jobs from an SQS standard queue and take up to five minutes to process each message. The queue uses the default 30-second visibility timeout. Monitoring shows many messages are processed two or three times even though no worker has crashed. What should a solutions architect do to stop the duplicate processing?

Choose one.

  • a
    Convert the queue to a FIFO queue to eliminate duplicates.

    FIFO deduplication targets duplicate sends within the deduplication window; a message whose visibility timeout expires mid-processing is legitimately redelivered in FIFO queues too.

  • b
    Enable long polling on the receive calls.

    Long polling reduces empty responses and API cost when the queue is quiet; it does not change how long a received message stays invisible.

  • c
    Increase the queue's visibility timeout to comfortably exceed the maximum processing time. Correct

    Messages are reappearing because the 30-second invisibility window expires while a healthy worker is still mid-job; extending the timeout past five minutes keeps in-flight messages hidden until processing finishes.

  • d
    Attach a dead-letter queue with a redrive policy.

    A DLQ isolates messages that repeatedly fail processing; these messages succeed, they are just redelivered mid-flight, so a DLQ would not fire and would not fix the cause.

The concept

SQS does not delete a message on receive; it hides it for the visibility timeout while the consumer works, and the consumer deletes it after success. A timeout shorter than processing time makes healthy in-flight messages reappear and be processed again.

Why that’s the answer

The symptom signature is exact: duplicates without worker crashes means the visibility timeout is expiring before processing completes, so the message becomes visible again and another worker picks it up. Raising the timeout above the maximum processing time (with margin) removes the premature reappearance and the duplicates with it. The FIFO conversion is the classic trap: FIFO's exactly-once guarantee deduplicates producer-side sends, but visibility-timeout expiry is a consumer-side redelivery that FIFO performs as well, so the problem would persist with a lower throughput cap added. Long polling addresses a completely different concern, the cost and latency of polling an empty queue. A dead-letter queue catches poison messages that keep failing; here every processing attempt succeeds, so the redrive counter is irrelevant to the duplication.

How to reason it out
  1. Note the two facts: processing takes up to five minutes and visibility timeout is 30 seconds.
  2. Recall the mechanism: an in-flight message reappears when its visibility timeout expires before the consumer deletes it.
  3. Conclude the duplicates are premature redeliveries, not producer duplicates or failures.
  4. Fix the mismatch by raising the visibility timeout above the worst-case processing time rather than changing queue type or adding a DLQ.

Exam tip: Messages processed twice by healthy workers means the visibility timeout is shorter than processing time; raise the timeout, do not switch to FIFO.

Scalable, Loosely Coupled Architectures: SQS, SNS, EventBridge, Step Functions — the lesson that teaches this.

Question 5Design Resilient Architectures

A consumer fleet processes messages from an SQS standard queue. A small number of malformed messages always throw an exception, return to the queue after the visibility timeout, and are redelivered indefinitely, wasting worker capacity and delaying valid messages. The team wants the bad messages isolated automatically for later analysis with the LEAST operational overhead. What should the architect configure?

Choose one.

  • a
    A redrive policy on the queue that moves messages to a dead-letter queue after a maximum receive count, with an alarm on the dead-letter queue depth. Correct

    After the configured number of failed receives, SQS automatically sidelines the poison message into the DLQ where it can be inspected, and the main queue keeps flowing.

  • b
    Add exception handling that deletes any message that fails processing.

    Deleting on first failure discards the evidence and also destroys valid messages that failed transiently, and it requires code changes rather than a queue setting.

  • c
    Purge the queue whenever the backlog of failing messages grows.

    Purging deletes every message in the queue, including all valid pending work, and requires a human to notice and act each time.

  • d
    Raise the visibility timeout so failing messages reappear less frequently.

    A longer timeout only stretches the interval between failed attempts; the poison messages still cycle forever and still consume worker attempts.

The concept

A dead-letter queue with a redrive policy is the built-in mechanism for poison messages: after maxReceiveCount failed receives, SQS moves the message aside automatically so it can be analyzed without blocking the main queue.

Why that’s the answer

The redrive policy is a queue configuration, not code, and it does exactly what the requirement states: isolate repeatedly failing messages automatically and preserve them for analysis, with an alarm providing visibility. Deleting on exception is the tempting code fix, but it destroys the malformed messages the team wants to analyze, cannot distinguish a poison message from a transient failure, and pushes failure policy into every consumer. Purging is a destructive manual action that sacrifices all valid in-flight work to remove a few bad messages. Raising the visibility timeout misreads the mechanism: the messages are not duplicating prematurely, they are genuinely failing, so slowing the retry loop just delays the same infinite cycle.

How to reason it out
  1. Identify the pattern: messages that always fail and recycle forever are poison messages.
  2. Recall the built-in answer: a redrive policy with maxReceiveCount moves them to a DLQ automatically.
  3. Verify the requirement fit: DLQ preserves the messages for analysis and unblocks the main queue with zero ongoing human effort.
  4. Eliminate destructive options (delete-on-failure, purge) because they lose data, and timeout tuning because the messages fail rather than duplicate.

Exam tip: Messages that repeatedly fail belong in a dead-letter queue via a redrive policy; never delete or purge your way around a poison message.

Scalable, Loosely Coupled Architectures: SQS, SNS, EventBridge, Step Functions — the lesson that teaches this.

Question 6Design Resilient Architectures

When a customer places an order, three systems owned by different teams must each receive and process every order event independently: inventory, billing, and an analytics pipeline. Each system must be able to work through its own backlog after brief downtime, and the company wants to add future consumers without changing the ordering application. Which architecture meets these requirements?

Choose one.

  • a
    Send each order event to a single SQS queue that all three systems poll.

    Consumers on one queue compete for messages, so each event is processed by only one of the three systems instead of all of them.

  • b
    Have the ordering application write each event to three separate SQS queues, one per system.

    This works today, but the publisher must implement three sends with partial-failure handling, and every new consumer requires changing and redeploying the ordering application.

  • c
    Publish each order event once to an SNS topic, with a separate SQS queue subscribed to the topic for each consuming system. Correct

    SNS fans a copy of every event to each subscribed queue, each queue buffers durably through consumer downtime, and a new consumer is just a new subscription with no publisher change.

  • d
    Have the ordering application call each system's REST API synchronously when an order is placed.

    Synchronous calls couple order placement to every consumer's availability; one slow or offline system delays or fails the order path and events are lost during downtime.

The concept

One event consumed independently by many systems is the fan-out pattern: an SNS topic delivers a copy to every subscriber, and subscribing an SQS queue per consumer adds durable buffering, independent retries, and independent scaling.

Why that’s the answer

SNS-to-SQS fan-out satisfies all three requirements at once: every subscriber receives its own copy (unlike a shared queue), each queue holds the backlog while its consumer is down (unlike direct push or synchronous calls), and adding a consumer is a subscription change that never touches the publisher. The shared single queue is the classic elimination: competing consumers each steal messages, so inventory would see only a fraction of the events. The three-queue publisher option is the strongest distractor because it functions, but it moves fan-out logic into application code, forcing the publisher to handle partial failures across three sends and to change whenever a consumer is added, which violates the stated extensibility requirement. Direct synchronous calls are tight coupling and lose events outright when a consumer is offline.

How to reason it out
  1. Identify the delivery semantics needed: every consumer must get every event, which means pub/sub, not competing consumption.
  2. Add durability per consumer: subscribe an SQS queue for each system so downtime becomes a backlog instead of loss.
  3. Check extensibility: new consumers are new subscriptions, leaving the ordering application untouched.
  4. Eliminate the shared queue (message stealing), publisher-managed multi-queue writes (coupling and code change per consumer), and synchronous calls (availability coupling).

Exam tip: Every-consumer-gets-a-copy plus per-consumer durability is SNS fanning out into one SQS queue per subscriber.

Scalable, Loosely Coupled Architectures: SQS, SNS, EventBridge, Step Functions — the lesson that teaches this.

What SAA-C03 domain 2 tests, topic by topic

The official exam guide breaks Design Resilient Architectures into 2 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published SAA-C03 practice questions per topic in Design Resilient Architectures
TopicWhat it coversQuestions
Design scalable and loosely coupled architecturesExam guide task 2.1. Event-driven, microservice, and multi-tier designs; decoupling with SQS and pub/sub messaging, API Gateway, and Step Functions orchestration; when to choose serverless (Lambda, Fargate) vs containers (ECS, EKS) and migrating apps into containers; horizontal vs vertical scaling, load balancing (ALB), caching, edge accelerators (CDN), read replicas, and storage types (object, file, block).20
Design highly available and/or fault-tolerant architecturesExam guide task 2.2. Multi-AZ and multi-Region design with Regions, AZs, and Route 53; disaster-recovery strategies (backup and restore, pilot light, warm standby, active-active failover) against RTO/RPO; mitigating single points of failure; failover and distributed design patterns, immutable infrastructure, RDS Proxy; data durability strategies, service quotas and throttling, workload visibility with X-Ray, and improving the reliability of legacy applications.20
Total40

Revise Design Resilient Architectures before you drill it

Other SAA-C03 domains

Design Resilient Architectures: your questions

Design Resilient Architectures is domain 2 of the SAA-C03 exam guide and carries 26% of the scored content — the 2nd-heaviest of the 4 domains. On a 65-question paper that works out to roughly 17 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official SAA-C03 exam guide.