A data engineering team runs a nightly job on EC2 that sequentially scans multi-terabyte log files stored on an attached EBS volume, computes aggregates, and writes summary output. The job needs high sequential throughput but performs almost no random I/O. Which volume type meets the throughput requirement MOST cost-effectively?
Choose one.
HDD volumes sell throughput and SSD volumes sell IOPS; for large sequential scans, st1 delivers the required throughput at HDD prices.
The access pattern is large and sequential with almost no random I/O — the exact workload st1 was built for, so it meets the performance requirement at the lowest cost that still performs. gp3 and io2 both work but pay SSD premiums for random-I/O capability that goes unused, which the MOST cost-effective qualifier punishes. sc1 is the trap in the other direction: it is the cheapest EBS volume, but it targets cold, rarely touched data and delivers lower throughput, so a nightly performance-sensitive scan can fall short of the requirement. Cheapest-that-still-meets-the-requirement is the rule, and that is st1 here.
- Identify the access pattern: large sequential reads, minimal random I/O.
- Map pattern to media: sequential throughput favors HDD volumes over SSD on price.
- Check the access frequency: nightly, throughput-sensitive scans are regular work, ruling out cold-tier sc1.
- Confirm st1 meets the throughput requirement at lower cost than gp3 or io2.
Exam tip: Sequential big-data scans resolve to st1 — SSDs overpay for unused IOPS, and sc1 undershoots on regularly accessed data.
High-Performing Storage: S3 vs EBS vs EFS, Volume Types, and Hybrid Options — the lesson that teaches this.