A retailer keeps 2 TB of order history in an Amazon DynamoDB table that serves production traffic with provisioned capacity. Each week an ML engineer needs a full copy of the table in Amazon S3 to build a training dataset. The copy must not consume the table's read capacity or require custom code to page through items. What should the engineer do?
Choose one.
DynamoDB export to Amazon S3 builds full or incremental exports from point-in-time recovery data without touching table capacity.
DynamoDB's export to S3 feature requires point-in-time recovery to be enabled and writes the table's data to an S3 prefix in DynamoDB JSON or Amazon Ion format. Because it reads from the continuous backup, it uses no read capacity and needs no paging code. The export has no built-in schedule, so a scheduler such as EventBridge Scheduler starts it each week. A Glue scan job also lands the data in S3, but it reads the live table and competes with production. Streams miss existing items, and a restored backup is another table rather than S3 data.
- Note the constraints: full copy in S3, no read capacity, no paging code.
- Recall that table exports read from point-in-time recovery backups.
- Eliminate scans (consume capacity) and streams (miss existing items).
- Enable PITR and start an export to S3 each week from a scheduler.
Exam tip: Full DynamoDB copy in S3 without read capacity: PITR plus export to S3.
Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.