SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
SOA-C03 · Domain 1

Monitoring, Logging, Analysis, Remediation, and Performance Optimization practice questions

Monitoring, Logging, Analysis, Remediation, and Performance Optimization is worth 22% of the SOA-C03 exam — the heaviest of the 5 domains. CloudWatch-centered observability — metrics, alarms, and logs — plus automated remediation and performance tuning for compute, storage, and databases. 6 fully worked examples are further down this page, answers included.

Exam weight
22%
the heaviest of the 5 domains
Questions
60
across 3 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 5 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Monitoring, Logging, Analysis, Remediation, and Performance Optimization questions, fully explained

Questions from the SOA-C03 bank mapped to domain 1, with the answer key and the reasoning behind every option. None of them repeat the examples on the main SOA-C03 practice page.

Question 1Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A CloudOps team must collect memory and disk-space metrics from 300 EC2 instances and keep the collection configuration identical across the entire fleet. Which approach meets the requirement with the LEAST operational overhead?

Choose one.

  • a
    Store the CloudWatch agent configuration in Systems Manager Parameter Store and use Systems Manager to install and configure the agent across the fleet. Correct

    One configuration in Parameter Store gives every instance an identical setup, and Systems Manager installs and configures the agent at scale without logging in to any instance.

  • b
    Connect to each instance over SSH, install the agent, and copy a local configuration file.

    Hand-configuring 300 instances is slow, requires SSH access, and local config files drift apart over time — the opposite of a consistent fleet.

  • c
    Enable detailed monitoring on all instances from the console.

    Detailed monitoring increases metric frequency to 1 minute but never adds memory or disk-space metrics, which only the agent can collect.

  • d
    Deploy a custom script to each instance that calls PutMetricData with memory values.

    Custom scripts duplicate what the agent already does, and 300 per-instance copies of homegrown code invite drift and maintenance burden.

The concept

The CloudWatch agent reads a JSON configuration describing which metrics and log files to collect. The operational best practice for fleets is to store that configuration centrally in Systems Manager Parameter Store and to install and configure the agent through Systems Manager rather than by hand.

Why that’s the answer

Parameter Store gives the whole fleet one authoritative agent configuration, and Systems Manager performs the installation and configuration at scale — no SSH, no per-instance drift. SSHing to each instance is unmanageable at 300 instances; detailed monitoring cannot produce in-guest metrics at all; and a custom PutMetricData script reinvents the agent with code the team must now maintain.

How to reason it out
  1. Author the CloudWatch agent JSON configuration listing the memory and disk metrics to collect.
  2. Store the configuration as a parameter in Systems Manager Parameter Store.
  3. Use Systems Manager to install the agent across the fleet.
  4. Configure the agent on every instance to load the shared parameter, then verify metrics arrive.

Exam tip: Fleet-wide agent rollout means Systems Manager for installation and Parameter Store for one shared agent configuration.

CloudWatch Metrics, Alarms, and Log Filters for SOA-C03 — the lesson that teaches this.

Question 2Monitoring, Logging, Analysis, Remediation, and Performance Optimization

An Auto Scaling group scales on average CPUUtilization, but scaling reacts slowly because the EC2 instance metrics arrive only every 5 minutes. What is the MOST direct way to make the instance metrics available at 1-minute frequency?

Choose one.

  • a
    Install the CloudWatch agent on the instances.

    The agent's job is collecting in-guest metrics such as memory and log files. The frequency of the standard CPUUtilization metric is controlled by detailed monitoring, not the agent.

  • b
    Publish CPU usage as a high-resolution custom metric with PutMetricData.

    This would require writing and maintaining custom collection code when a built-in setting already delivers 1-minute CPUUtilization.

  • c
    Enable detailed monitoring on the instances. Correct

    Detailed monitoring raises EC2 instance metrics from the free 5-minute standard frequency to 1-minute frequency for an additional cost — exactly what faster scaling reactions need.

  • d
    Shorten the alarm period on the scaling policy to 1 minute.

    An alarm cannot evaluate faster than datapoints arrive. With 5-minute data, 1-minute periods just produce missing datapoints.

The concept

EC2 standard monitoring publishes instance metrics every 5 minutes at no charge. Detailed monitoring publishes the same metrics every 1 minute for an additional cost — it changes frequency only, adding no new metrics — and is enabled per instance or in a launch template.

Why that’s the answer

Enabling detailed monitoring is the built-in, direct way to get 1-minute EC2 metrics so alarms and Auto Scaling react faster. The CloudWatch agent solves a different problem (in-guest metrics and logs); a high-resolution custom metric means writing collection code for data AWS already provides; and shrinking the alarm period cannot conjure datapoints that are only published every 5 minutes.

How to reason it out
  1. Identify that the bottleneck is metric frequency, not metric availability.
  2. Enable detailed monitoring on the instances or in their launch template.
  3. Confirm CPUUtilization now arrives at 1-minute granularity.
  4. Adjust the scaling policy's alarm period to evaluate 1-minute datapoints.

Exam tip: Standard monitoring is 5-minute and free; detailed monitoring is 1-minute and paid — and it changes frequency, never the metric set.

CloudWatch Metrics, Alarms, and Log Filters for SOA-C03 — the lesson that teaches this.

Question 3Monitoring, Logging, Analysis, Remediation, and Performance Optimization

An application's internal queue depth can spike and cause failures within about 30 seconds. The team publishes the queue depth to CloudWatch with PutMetricData and needs an alarm that can react in well under a minute. Which combination of settings makes this possible?

Choose one.

  • a
    Publish the metric at standard resolution and create an alarm with a 1-minute period.

    Standard-resolution custom metrics store data at 1-minute granularity, so the fastest evaluation is a full minute — too slow for a 30-second failure.

  • b
    Enable detailed monitoring for the application's instances.

    Detailed monitoring applies to EC2 instance metrics and delivers 1-minute frequency. It has no effect on a custom application metric and is still too slow.

  • c
    Create a metric filter on the application log group with a 10-second alarm period.

    This adds a logging pipeline the queue depth does not need, and it does not provide the 1-second high-resolution publication that a 10-second alarm period requires.

  • d
    Publish the metric as a high-resolution metric with 1-second storage resolution, and create a high-resolution alarm with a 10-second period. Correct

    High-resolution custom metrics store data down to 1-second granularity, and high-resolution alarms can evaluate on 10-second or 30-second periods — fast enough to catch a 30-second failure window.

The concept

Custom metrics come in two resolutions: standard resolution stores data at 1-minute granularity, while high-resolution metrics store data down to 1-second granularity. Only high-resolution metrics support high-resolution alarms, which evaluate on 10-second or 30-second periods.

Why that’s the answer

Reacting inside a 30-second window requires both halves: publishing the metric at high resolution so datapoints exist at second-level granularity, and a high-resolution alarm with a 10-second period to evaluate them. Standard resolution caps evaluation at 1-minute periods; detailed monitoring governs EC2 instance metrics, not custom metrics; and a metric filter is a log-pattern tool that does not deliver 1-second data.

How to reason it out
  1. Decide the required detection speed — here, well under a minute.
  2. Publish the queue-depth metric with PutMetricData as a high-resolution metric.
  3. Create a high-resolution alarm with a 10-second period on the metric.
  4. Attach a notification or remediation action and test with a simulated spike.

Exam tip: Sub-minute detection needs a high-resolution custom metric plus a high-resolution alarm on a 10- or 30-second period.

CloudWatch Metrics, Alarms, and Log Filters for SOA-C03 — the lesson that teaches this.

Question 4Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A team runs microservices on an Amazon ECS cluster and wants task-level and service-level CPU and memory metrics, plus container logs, available in CloudWatch with the LEAST setup effort. Which option should the team choose?

Choose one.

  • a
    Manually install the CloudWatch agent on each container instance and hand-build per-task metric dimensions.

    This reimplements what Container Insights already packages, with far more setup and ongoing maintenance.

  • b
    Migrate the cluster's monitoring to Amazon Managed Service for Prometheus.

    Managed Prometheus suits teams standardized on Prometheus-format metrics; the requirement here is CloudWatch-native metrics and logs with minimal effort.

  • c
    Create subscription filters on every task's log stream.

    Subscription filters deliver log events to destinations such as Kinesis or Lambda; they do not produce CPU or memory metrics for tasks and services.

  • d
    Enable Container Insights for the cluster. Correct

    Container Insights is the packaged CloudWatch capability that collects cluster-, service-, and task-level metrics and logs from ECS with minimal configuration.

The concept

Container Insights deploys the collection machinery into Amazon ECS or EKS clusters to gather cluster-, service-, task-, and pod-level metrics and logs into CloudWatch — the purpose-built path for container observability in CloudWatch.

Why that’s the answer

Enabling Container Insights delivers exactly the requested task- and service-level metrics plus container logs with the least effort, because it is the packaged feature for this job. Hand-installing agents and inventing dimensions rebuilds the same capability manually; Managed Prometheus answers a different requirement (Prometheus-standardized teams); and subscription filters move log events rather than producing container metrics.

How to reason it out
  1. Enable Container Insights on the ECS cluster.
  2. Confirm task- and service-level CPU and memory metrics appear in CloudWatch.
  3. Verify container logs are flowing into the cluster's log groups.
  4. Build alarms and dashboards on the collected metrics.

Exam tip: For ECS or EKS metrics and logs in CloudWatch, enable Container Insights — do not hand-roll agent deployments.

CloudWatch Metrics, Alarms, and Log Filters for SOA-C03 — the lesson that teaches this.

Question 5Monitoring, Logging, Analysis, Remediation, and Performance Optimization

An application writes lines containing the word "FATAL" to a log file that the CloudWatch agent already ships to a log group. The on-call team must be paged when five or more FATAL lines occur within 5 minutes. Which solution requires the LEAST operational overhead?

Choose one.

  • a
    Create a subscription filter that streams matching events to a Lambda function that counts them and publishes to SNS.

    This works but requires writing, testing, and maintaining custom counting code that a metric filter plus alarm already provides natively.

  • b
    Run a CloudWatch Logs Insights query for FATAL lines whenever the team suspects a problem.

    Logs Insights is an on-demand query engine; it does not continuously watch logs or page anyone automatically.

  • c
    Enable CloudTrail data events for the application's log group.

    CloudTrail records API calls, not application log content — it cannot detect FATAL lines inside log events.

  • d
    Create a metric filter matching FATAL that publishes a count metric, then create an alarm on that metric with a sum threshold of 5 over 5 minutes and an SNS notification action. Correct

    This is the built-in log-error-alarm pipeline: pattern match becomes a metric, the metric drives an alarm, and the alarm notifies SNS — no code at all.

The concept

A metric filter watches a log group for a pattern and increments a CloudWatch metric on every match, converting unstructured log lines into an alarmable number. Pairing it with an alarm and an SNS action is the standard zero-code log-alerting pipeline.

Why that’s the answer

The metric filter to alarm to SNS chain meets the requirement entirely with managed configuration: pattern FATAL, count metric, alarm on sum of at least 5 over 5 minutes, SNS page. The subscription-filter-plus-Lambda option achieves the same result but adds custom code the team must maintain — failing the least-overhead constraint. Logs Insights only answers questions when someone asks them, and CloudTrail audits API activity, not application log text.

How to reason it out
  1. Create a metric filter on the log group with the pattern FATAL, publishing a count metric with value 1 per match.
  2. Set a default value of 0 so the metric emits datapoints during quiet periods.
  3. Create an alarm: sum of the metric at least 5 over a 5-minute period.
  4. Set the alarm action to an SNS topic with a confirmed on-call subscription.
  5. Test by writing FATAL lines and confirming the page arrives.

Exam tip: Alert on a log pattern with a metric filter feeding an alarm — reserve subscription filters and Lambda for when the log data itself must flow somewhere.

CloudWatch Metrics, Alarms, and Log Filters for SOA-C03 — the lesson that teaches this.

Question 6Monitoring, Logging, Analysis, Remediation, and Performance Optimization

After an incident, an engineer creates a metric filter to count occurrences of a specific error string in a log group and expects the new metric to show yesterday's error spike. The metric has no datapoints for any past period, although the log events are still visible in the log group. What explains this behavior?

Choose one.

  • a
    The log group's retention policy deleted yesterday's events before the filter ran.

    The scenario states the log events are still visible in the group, so retention did not remove them — and even present events are not evaluated retroactively.

  • b
    Metric filters take 24 hours before any datapoints appear.

    There is no 24-hour warm-up; the metric receives datapoints as soon as new matching events are ingested.

  • c
    A default value must be configured before the metric filter can emit any datapoints.

    A default value only makes the metric emit 0 during quiet periods going forward; it has nothing to do with historical data.

  • d
    Metric filters only evaluate log events ingested after the filter is created; they never backfill historical data. Correct

    This is a documented property of metric filters: they apply to new log data only, so yesterday's events produce no metric datapoints.

The concept

Metric filters apply only to log data ingested after the filter is created. They convert future pattern matches into metric datapoints and never re-scan the log group's history.

Why that’s the answer

The absence of historical datapoints is the designed behavior — no backfill exists. Retention is ruled out because the events remain visible (and retention would not enable retroactive evaluation anyway); there is no 24-hour delay before metric filters work; and a default value shapes quiet-period behavior for new data, not history. To analyze the past spike, the engineer should query the existing events with CloudWatch Logs Insights instead.

How to reason it out
  1. Create the metric filter to cover future occurrences of the error pattern.
  2. For the historical spike, run a Logs Insights query over yesterday's time range to count matches.
  3. Add a default value of 0 to the filter so future quiet periods still emit datapoints.
  4. Attach an alarm to the new metric for the next occurrence.

Exam tip: Metric filters never backfill — analyze historical log data with Logs Insights, and use the filter for everything ingested from now on.

CloudWatch Metrics, Alarms, and Log Filters for SOA-C03 — the lesson that teaches this.

What SOA-C03 domain 1 tests, topic by topic

The official exam guide breaks Monitoring, Logging, Analysis, Remediation, and Performance Optimization into 3 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published SOA-C03 practice questions per topic in Monitoring, Logging, Analysis, Remediation, and Performance Optimization
TopicWhat it coversQuestions
Implement metrics, alarms, and filters by using AWS monitoring and logging servicesExam guide task 1.1. Monitoring and logging for serverless, compute, and AI workloads with CloudWatch, CloudTrail, and Amazon Managed Service for Prometheus; configuring and managing the CloudWatch agent for EC2, ECS, and EKS metrics/logs; CloudWatch alarms (including composite alarms) that invoke services directly or via EventBridge, with SNS notifications; shareable cross-account, cross-Region dashboards.20
Identify and remediate issues by using monitoring and availability metricsExam guide task 1.2. Analyzing performance metrics and automating remediation with CloudWatch, Lambda, Systems Manager, CloudTrail, Kiro, and AWS DevOps Agent; routing, enriching, and delivering events with EventBridge and troubleshooting event bus rules; creating and running custom and predefined Systems Manager Automation runbooks.20
Implement performance optimization strategies for compute, storage, and database resourcesExam guide task 1.3. Optimizing compute via performance metrics, resource tags, and AWS tools (EC2 instance selection, placement groups); analyzing and optimizing EBS performance and volume types; S3 performance strategies (DataSync, Transfer Acceleration, multipart uploads, lifecycle policies); selecting and optimizing shared storage (EFS, FSx, Amazon S3 Files, EFS lifecycle policies); monitoring RDS with Performance Insights and CloudWatch alarms and tuning with proactive recommendations and RDS Proxy.20
Total60

Revise Monitoring, Logging, Analysis, Remediation, and Performance Optimization before you drill it

Other SOA-C03 domains

Monitoring, Logging, Analysis, Remediation, and Performance Optimization: your questions

Monitoring, Logging, Analysis, Remediation, and Performance Optimization is domain 1 of the SOA-C03 exam guide and carries 22% of the scored content — the heaviest of the 5 domains. On a 65-question paper that works out to roughly 14 questions, though AWS does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official SOA-C03 exam guide.