A chat application streams answers from Amazon Bedrock by using the ConverseStream API. After a prompt template change, users complain that nothing appears on screen for several seconds after they press Enter, although the complete answer finishes in about the same total time as before. The ML engineer wants a CloudWatch alarm that tracks exactly this experience. Which metric should the alarm use?
Choose one.
For streaming Bedrock calls, TimeToFirstToken measures the wait before output starts; InvocationLatency measures until the last token.
The users' complaint is the delay before the first text appears, while overall duration is unchanged. TimeToFirstToken is emitted for ConverseStream and InvokeModelWithResponseStream and measures exactly that gap. InvocationLatency covers the whole response and would hide the regression; token counts and quota estimates describe volume, not responsiveness.
- Separate time-to-first-output from total response time.
- Note the call is streaming, so a first-token metric exists.
- Pick the metric that changed in the users' description.
Exam tip: Streaming responsiveness = TimeToFirstToken; total response time = InvocationLatency.
Monitoring ML Models, FMs and Agents in Production: Drift, A/B Tests and Bedrock Evaluations (MLA-C02) — the lesson that teaches this.