A radiology startup runs a GPU-based segmentation model in SageMaker AI. Each request is a 300 MB scan, and inference takes 6 to 9 minutes per scan. Scans arrive unpredictably, sometimes none for a whole day. Clinicians need results within 20 minutes of upload and want a notification when a result is ready. The startup does not want to pay for GPU instances while no scans are waiting. Which deployment meets these requirements?
Choose one.
Asynchronous inference queues large, long-running requests and can scale to zero.
The deciding facts combine: payload size (300 MB) and processing time (minutes) rule out real-time and serverless inference; the need to pay nothing while idle rules out an always-on endpoint; and the 20-minute deadline rules out a scheduled batch job. Asynchronous inference accepts requests via S3 up to 1 GB, processes for up to an hour, scales to zero instances when nothing is queued and notifies through SNS.
- Compare the payload and processing time with each option's limits.
- Eliminate options that cannot use a GPU or that keep an instance running.
- Check the turnaround deadline against scheduled batch jobs.
- Choose asynchronous inference with scale-to-zero and SNS notifications.
Exam tip: Large payloads, minutes per request, near-real-time results, idle periods: asynchronous inference.
Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02) — the lesson that teaches this.