An ML engineer trains a PyTorch model on 3 TB of TFRecord shards stored under one Amazon S3 prefix. Each SageMaker AI training job uses the default input mode and spends about 40 minutes copying data before the first epoch begins. The training script opens the shards through ordinary file paths, and the team does not want to change the script or provision new storage. What should the engineer do?
Choose one.
SageMaker AI training input modes: File (copy first), FastFile (stream with file semantics) and Pipe (stream through FIFO pipes).
File mode, the default, copies the whole dataset to the instance's storage before the script runs, which is where the 40 minutes go. FastFile mode keeps POSIX file access but streams objects from S3 as the script reads them, so the job starts almost immediately and needs no code change. Pipe mode also streams but changes the interface to named pipes, and FSx for Lustre works but adds a file system and VPC setup the team does not want.
- Identify the cause: File mode downloads the full dataset before training starts.
- Note the constraints: no script change and no new storage service.
- Rule out Pipe mode (different read interface) and FSx for Lustre (new infrastructure).
- Choose FastFile mode on the same S3 prefix.
Exam tip: Slow start in File mode with a file-based script and S3 data: switch to FastFile.
Collecting and Storing Data for ML and AI on AWS (MLA-C02) — the lesson that teaches this.