SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
SOA-C03 · Domain 2

Reliability and Business Continuity practice questions

Reliability and Business Continuity is worth 22% of the SOA-C03 exam — the heaviest of the 5 domains. Scalability and elasticity, highly available and resilient environments, and backup/restore strategies that meet RTO and RPO. 6 fully worked examples are further down this page, answers included.

Exam weight
22%
the heaviest of the 5 domains
Questions
60
across 3 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 5 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Reliability and Business Continuity questions, fully explained

Questions from the SOA-C03 bank mapped to domain 2, with the answer key and the reasoning behind every option. None of them repeat the examples on the main SOA-C03 practice page.

Question 1Reliability and Business Continuity

An operations team wants an Auto Scaling group to add one instance when average CPU exceeds 70% but add four instances at once when average CPU exceeds 90%. Which scaling policy type supports this tiered response?

Choose one.

  • a
    Target tracking scaling

    Target tracking takes a single target value and manages its own alarms; it does not let you define different adjustment sizes for different breach severities.

  • b
    Step scaling Correct

    Step scaling defines step adjustments tied to alarm breach magnitude — +1 instance for the 70–90% band and +4 above 90% — so the response scales with severity.

  • c
    Simple scaling

    Simple scaling ties one alarm to one fixed adjustment and then waits out a cooldown, so it cannot respond differently to a small breach versus a large one.

  • d
    Scheduled scaling

    Scheduled scaling changes capacity at specific times on a schedule; it does not react to CPU at all.

The concept

Step scaling sizes the capacity adjustment to how far the metric breached the alarm threshold, using step adjustments you define.

Why that’s the answer

Only step scaling lets you say 'a small breach adds one instance, a large breach adds four' — the adjustment scales with breach magnitude. Target tracking holds one value and hides its alarms; simple scaling makes exactly one fixed adjustment per cooldown; scheduled scaling ignores metrics entirely.

How to reason it out
  1. Identify the requirement: different response sizes at different severity levels of the same metric.
  2. Create a CloudWatch alarm on average CPU for the Auto Scaling group.
  3. Attach a step scaling policy with step adjustments: +1 instance for the 70–90% band, +4 instances above 90%.
  4. Configure instance warmup so newly launched instances are counted correctly during multi-step responses.

Exam tip: Tiered responses sized to breach severity are the signature use case for step scaling.

Auto Scaling and Elasticity: EC2, Caching, and Database Scaling — the lesson that teaches this.

Question 2Reliability and Business Continuity

A payroll application's EC2 fleet is overwhelmed every weekday at 9:00 AM when employees log in, and traffic falls off sharply after 6:00 PM. The pattern is identical every business day. Which scaling approach ensures capacity is already in place before the morning surge?

Choose one.

  • a
    A simple scaling policy triggered by a CPU alarm

    Simple scaling is reactive and legacy — instances launch only after CPU has already spiked, so the 9:00 AM users experience the slowdown before capacity arrives.

  • b
    Raising the Auto Scaling group's maximum capacity

    A higher maximum only permits more instances; it does not launch anything by itself, so the morning surge still hits an undersized fleet.

  • c
    Scheduled scaling actions that raise desired capacity before 9:00 AM on weekdays and lower it in the evening Correct

    Scheduled scaling sets min, desired, and max at specific recurring times, so the fleet is grown before the known surge instead of reacting after users are already waiting.

  • d
    Shortening the health-check grace period so instances enter service faster

    The grace period controls when health evaluation begins on new instances; it does not decide when or whether capacity is launched.

The concept

Scheduled scaling changes an Auto Scaling group's min, desired, and max capacity at specific times or on a recurring schedule, which suits predictable demand cycles.

Why that’s the answer

The load pattern is known and repeats every business day, so capacity should be provisioned ahead of time rather than in reaction to a breaching metric. Reactive policies (simple, step, target tracking) all launch after load has risen; raising the max alone launches nothing; the grace period is a health-evaluation setting, not a capacity trigger.

How to reason it out
  1. Confirm the demand pattern is time-driven and recurring, not random.
  2. Create a recurring scheduled action that raises min and desired capacity shortly before 9:00 AM on weekdays.
  3. Create a second recurring action that lowers capacity after 6:00 PM.
  4. Optionally keep a target tracking policy alongside for unexpected intraday spikes.

Exam tip: Predictable, clock-driven demand cycles are the trigger phrase for scheduled scaling — provision before the surge, not after.

Auto Scaling and Elasticity: EC2, Caching, and Database Scaling — the lesson that teaches this.

Question 3Reliability and Business Continuity

An e-commerce fleet shows a recurring daily traffic curve, and its application servers take nearly 15 minutes to bootstrap before they can serve requests. Reactive scaling policies consistently add capacity too late. Which policy type launches instances ahead of forecasted demand?

Choose one.

  • a
    Target tracking scaling

    Target tracking is reactive: it scales after the metric moves, so 15-minute bootstraps still leave a window where demand outruns capacity.

  • b
    Step scaling

    Step scaling also reacts to alarm breaches that have already happened; larger steps do not fix the lead-time problem.

  • c
    Simple scaling

    Simple scaling is the slowest reactive option — one fixed adjustment per cooldown window — making the late-capacity problem worse, not better.

  • d
    Predictive scaling Correct

    Predictive scaling uses a machine-learning forecast of historical load to launch capacity ahead of expected demand — exactly what slow-booting instances on a recurring pattern need.

The concept

Predictive scaling forecasts load from historical patterns with machine learning and launches capacity before the demand arrives.

Why that’s the answer

The scenario combines the two cues for predictive scaling: a recurring pattern (so the forecast has something to learn) and long instance lead time (so reactive launches arrive too late). All three reactive policies — target tracking, step, and simple — only act after the metric has moved, which is precisely the failure being described.

How to reason it out
  1. Recognize the two cues: recurring demand pattern plus long instance startup lead time.
  2. Enable predictive scaling on the Auto Scaling group so it learns the historical load curve.
  3. Keep a target tracking policy alongside to absorb deviations from the forecast.
  4. Validate that instances launched by the forecast are in service before the daily peak.

Exam tip: Recurring patterns plus instances that need lead time to be ready point to predictive scaling.

Auto Scaling and Elasticity: EC2, Caching, and Database Scaling — the lesson that teaches this.

Question 4Reliability and Business Continuity

A scale-out CloudWatch alarm for an EC2 Auto Scaling group has been in ALARM state for 30 minutes, but the group has not launched any instances and users report slow responses. Where should the engineer look FIRST to identify the cause?

Choose one.

  • a
    The health-check grace period on the Auto Scaling group

    The grace period delays health evaluation of instances that have already launched; it has no effect on whether the group launches instances at all.

  • b
    The deregistration delay on the load balancer target group

    Deregistration delay governs how long draining targets finish in-flight requests during scale-in or removal — it is unrelated to scale-out launches.

  • c
    The group's activity history, and whether desired capacity already equals the maximum Correct

    The activity history shows failed launches and their reasons, and desired-equals-max is the most common cause of an alarm that breaches with nothing launching — these are always the first two stops.

  • d
    Whether the group should be migrated from a launch template to a launch configuration

    Launch configurations are the older, immutable mechanism that no longer receives new features; migrating backward would not explain or fix missing launches.

The concept

When an alarm breaches but no instances launch, either the group has no headroom (desired already equals max) or launches are being attempted and failing — and the activity history records both.

Why that’s the answer

Reading the activity history and comparing desired to max covers the main causes in one check: capacity stuck at max, failed launches from an EC2 instance quota, an instance type unavailable in an AZ, a deleted launch-template AMI, or a suspended Launch process. Grace period and deregistration delay affect health evaluation and connection draining respectively, not launches; launch configurations are legacy and irrelevant here.

How to reason it out
  1. Open the Auto Scaling group's activity history and look for failed or absent scaling activities.
  2. Compare desired capacity to maximum — if they are equal, raise the maximum if the load is legitimate.
  3. If activities show launch failures, act on the stated reason: quota increase, alternate instance type or AZ, or a valid AMI.
  4. Check the suspended-processes list — a suspended Launch process silently stops all launches.

Exam tip: Alarm in breach with nothing launching → read the activity history first, and check whether desired already equals max.

Auto Scaling and Elasticity: EC2, Caching, and Database Scaling — the lesson that teaches this.

Question 5Reliability and Business Continuity

After a new deployment, an Auto Scaling group that uses ELB health checks terminates every newly launched instance about 90 seconds after launch — before the Java application finishes its three-minute startup. Instances that happen to survive serve traffic normally. What should the engineer change?

Choose one.

  • a
    Increase the health-check grace period beyond the application's startup time Correct

    The grace period is the window after launch during which health evaluation is deferred; setting it longer than the three-minute startup stops the group from killing instances that are still booting.

  • b
    Raise the unhealthy threshold on the load balancer target group

    A higher threshold only delays how fast the target group marks a target unhealthy; the designed control for slow startup is the grace period, which pauses the group's health evaluation entirely until the app can be up.

  • c
    Enable scale-in protection on new instances

    Scale-in protection shields instances from being selected during scale-in events; it does not prevent replacement of instances the group considers unhealthy.

  • d
    Shorten the health-check interval so instances are evaluated sooner

    Probing more often makes the problem worse — the still-booting application fails the checks even faster.

The concept

The health-check grace period gives a newly launched instance time to boot and start its application before the Auto Scaling group begins evaluating its health.

Why that’s the answer

Instances dying at ~90 seconds while the app needs three minutes is the textbook symptom of a grace period shorter than real startup time: the group evaluates health mid-boot, sees failures, and terminates. Raising the target group's unhealthy threshold nibbles at the symptom but the grace period is the purpose-built control; scale-in protection covers scale-in selection, not health replacement; faster probing accelerates the false failures.

How to reason it out
  1. Measure the application's true startup time from instance launch to serving traffic.
  2. Set the Auto Scaling group's health-check grace period comfortably above that time.
  3. Confirm new launches now survive and enter service.
  4. Distinguish this from instance warmup, which controls when a new instance's metrics join the group aggregate.

Exam tip: New instances terminated before the app boots → the health-check grace period is shorter than real startup time.

Auto Scaling and Elasticity: EC2, Caching, and Database Scaling — the lesson that teaches this.

Question 6Reliability and Business Continuity

An Auto Scaling group runs worker instances that process video-rendering jobs lasting up to two hours. When the group scales in during quiet periods, it sometimes terminates instances mid-job, forcing those jobs to restart. Which feature prevents busy instances from being selected for termination during scale-in?

Choose one.

  • a
    A Terminating:Wait lifecycle hook

    A lifecycle hook only pauses a termination that has already been decided, and the hook times out — it buys drain time, but a two-hour render on the selected instance still gets interrupted.

  • b
    The OldestInstance termination policy

    Termination policies change which instance is picked, not whether a busy one can be picked — the oldest instance may be mid-job too.

  • c
    A longer cooldown period on the scale-in policy

    Cooldown changes how often scale-in activities can occur; when one does occur, a busy instance can still be selected and terminated.

  • d
    Instance scale-in protection, set while an instance is processing a job and cleared when it is idle Correct

    Scale-in protection shields protected instances from being selected when the group scales in — the worker enables it when a job starts and removes it when idle, so only idle instances are terminated.

The concept

Scale-in protection — settable on the group or per instance — excludes instances from being selected for termination when the Auto Scaling group scales in.

Why that’s the answer

The requirement is that instances doing long-running work must never be chosen at scale-in, which is exactly what per-instance scale-in protection provides. A Terminating:Wait lifecycle hook is the tempting distractor: it delays a termination already in motion and expires on a timeout, so a two-hour job still loses. Termination policies reorder selection without protecting busy instances, and cooldown only spaces out scale-in events.

How to reason it out
  1. Have the worker set instance scale-in protection when it picks up a job.
  2. Have it clear the protection when the job completes, making the instance eligible again.
  3. Keep the scale-in policy as is — idle instances are still terminated normally.
  4. Optionally add a Terminating:Wait hook for graceful cleanup of the instances that are selected.

Exam tip: Long-running work surviving scale-in is the use case for instance scale-in protection, not lifecycle hooks.

Auto Scaling and Elasticity: EC2, Caching, and Database Scaling — the lesson that teaches this.

What SOA-C03 domain 2 tests, topic by topic

The official exam guide breaks Reliability and Business Continuity into 3 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published SOA-C03 practice questions per topic in Reliability and Business Continuity
TopicWhat it coversQuestions
Implement scalability and elasticityExam guide task 2.1. Configuring and managing scaling mechanisms across compute environments; caching for dynamic scalability (CloudFront, ElastiCache); scaling in managed databases (RDS, DynamoDB).20
Implement highly available and resilient environmentsExam guide task 2.2. Configuring and troubleshooting Elastic Load Balancing and Route 53 health checks; configuring fault-tolerant systems such as Multi-AZ deployments.20
Implement backup and restore strategiesExam guide task 2.3. Automating snapshots and backups for EC2, RDS, EBS, S3, and DynamoDB with services like AWS Backup; restoring databases (e.g. point-in-time restore) to meet RTO, RPO, and cost requirements; versioning for storage services (S3, FSx); disaster-recovery procedures and best practices (backup and restore, pilot light, warm standby, active/active).20
Total60

Revise Reliability and Business Continuity before you drill it

Other SOA-C03 domains

Reliability and Business Continuity: your questions

Reliability and Business Continuity is domain 2 of the SOA-C03 exam guide and carries 22% of the scored content — the heaviest of the 5 domains. On a 65-question paper that works out to roughly 14 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official SOA-C03 exam guide.