Optimising ML and GenAI Inference Cost and Performance on AWS: Instances, Savings Plans, Tokens and Vectors (MLA-C02)
Optimising ML and AI infrastructure cost and performance means running models, foundation models and agents on just enough of the right capacity, bought the cheapest way, while you can see where time and money go. For MLA-C02 Task 4.2 that covers choosing inference instance families, troubleshooting with Amazon CloudWatch, AWS X-Ray and Amazon Bedrock AgentCore Observability, building performance dashboards, rightsizing and scheduling endpoint capacity, tagging and capping spend with AWS cost tools, Savings Plans, and the AI-specific cost drivers: FM pricing options, agent runtime consumption, tokens, embeddings and vector storage.
On this page10 sections
- How do you choose an inference instance family?
- Which metrics and tools troubleshoot a slow endpoint?
- How do you build performance dashboards in CloudWatch?
- How do you optimise endpoint capacity for cost, performance and reliability?
- Which cost management tools track and cap ML and AI spend?
- Which purchasing options lower ML infrastructure cost?
- What drives the cost of foundation model inference in production?
- How do you monitor and reduce agent resource consumption?
- How do you control token usage and quota burn?
- How do you control embedding and vector storage costs?
- Choose an inference instance family and size from utilisation metrics and GPU memory, and use Inference Recommender default and advanced jobs to confirm it.
- Troubleshoot endpoint, request-path and agent latency with CloudWatch metrics, AWS X-Ray tracing and sampling, and AgentCore Observability.
- Build CloudWatch performance dashboards with percentile statistics, cross-account observability and SEARCH expressions.
- Optimise endpoint capacity for cost and reliability with multi-AZ instance counts, scheduled scaling and multi-model endpoints.
- Attribute, forecast and cap ML and AI spend with cost allocation tags, application inference profiles, AWS Budgets, Cost Anomaly Detection and Savings Plans.
- Compare FM inference pricing options (on-demand, Flex and Priority tiers, batch, reserved capacity, self-hosting) and manage agent runtime consumption.
- Reduce token, embedding and vector storage costs with output-token sizing, prompt caching, incremental sync, batch embedding and smaller vectors.
How do you choose an inference instance family?
Choose the instance family from the resource the model actually runs out of, not from the model's name or the instance you trained on. A SageMaker AI real-time endpoint pays for every vCPU, GiB of RAM and accelerator on the instance every hour, so any resource that sits idle is pure cost. Read the endpoint's utilisation metrics first, then move to the family that supplies the bottleneck resource and little else.
| What the metrics show | Family to consider | Why |
|---|---|---|
| CPUUtilization high, MemoryUtilization low, no GPU needed (scikit-learn, XGBoost, small NLP) | Compute optimised ml.c* | Same vCPUs as the general-purpose size with half the memory, at a lower hourly price |
| Balanced CPU and memory use | General purpose ml.m* | Roughly 4 GiB per vCPU; the sensible default |
| Light compute but large in-memory tables, embeddings or caches | Memory optimised ml.r* | Roughly 8 GiB per vCPU, so the RAM fits without paying for unused vCPUs |
| Deep learning or LLM inference that needs a GPU | Accelerated ml.g4dn, ml.g5, ml.g6, ml.g6e, ml.p* | Choose by GPU memory first, then GPU speed and price |
| High-volume deep learning inference on a supported model | AWS Inferentia2 ml.inf2 (or Graviton ml.*g for CPU work) | Lower cost per inference, after compiling with the Neuron SDK |
Two rules catch most mistakes. First, an accelerator the model does not use (a classical model on ml.g5 with GPUUtilization near zero) is always the first thing to remove. Second, adding a bigger instance of a family that already holds the bottleneck is usually not cheaper than moving to the right family: an ml.m5.4xlarge costs the same as two ml.m5.2xlarge, so consolidating changes nothing.
GPU memory decides LLM fit
For self-hosted LLMs the binding limit is GPU memory. Weights in FP16 or BF16 take about 2 bytes per parameter, so a 13-billion-parameter model needs about 26 GB for weights alone, plus headroom for the KV cache that grows with batch size and context length. Useful figures: NVIDIA T4 (g4dn) has 16 GB, A10G (g5) and L4 (g6) have 24 GB, and L40S (g6e) has 48 GB per GPU. The single-GPU sizes of a family all share the same GPU: moving from ml.g5.xlarge to ml.g5.8xlarge adds vCPUs and host RAM, not GPU memory. The larger multi-GPU sizes (such as ml.g5.12xlarge with four GPUs) fit bigger models only by splitting them across GPUs with tensor parallelism. When a model just overflows a 24 GB GPU, the smallest instance with a 48 GB GPU is usually cheaper than a multi-GPU size. Quantisation (8-bit or 4-bit weights) is the other lever, at some cost in quality.
The cost comparison between Inferentia, Graviton and GPU instances for a given deployment, and model compilation, are covered in the Task 3.1 lesson on deployment; this lesson focuses on rightsizing a running workload.
Confirm the choice with SageMaker Inference Recommender
Amazon SageMaker Inference Recommender load-tests a model on candidate instance types and reports latency, throughput and cost per inference for each, so you do not have to build a benchmarking harness. It works on a model package version in SageMaker Model Registry (or a model you describe with its container, artefact and sample payload).
| Default (instance recommendations) job | Advanced (load test) job | |
|---|---|---|
| Purpose | Quick shortlist of instance types for a model | Benchmark specific configurations against your real traffic and latency targets |
| You provide | The model package and a sample payload | The endpoint configurations (instance types, counts) to test, a traffic pattern with phases (initial users, spawn rate, duration) and stopping conditions |
| Stopping conditions | Chosen by the service | You set them, for example stop when P95 or P99 ModelLatency exceeds a threshold or when errors appear |
| Typical run time | Typically about 45 minutes | Longer, depending on the traffic phases |
Use a default job first, then an advanced job when the requirement specifies the traffic pattern to simulate or a latency target the test must enforce: only an advanced job takes traffic phases and stopping conditions. Benchmarking belongs before launch, not on production traffic. Inference Recommender is a pre-deployment tool: SageMaker Model Monitor watches drift after deployment, SageMaker Debugger profiles training jobs, and AWS Cost Explorer rightsizing recommendations cover EC2 instances, not SageMaker AI endpoint sizing.
Which metrics and tools troubleshoot a slow endpoint?
Start with the endpoint's own CloudWatch metrics to decide whether the time is spent inside SageMaker AI or somewhere else, then trace the rest of the request path with AWS X-Ray.
| Metric | Namespace | What it tells you |
|---|---|---|
ModelLatency | AWS/SageMaker | Time the model container took to respond, as seen by SageMaker AI |
OverheadLatency | AWS/SageMaker | Time SageMaker AI added on top (routing, authentication, payload handling) |
ModelSetupTime | AWS/SageMaker | Serverless endpoints only: time to launch compute and load the model, the cold-start cost |
CPUUtilization, MemoryUtilization | /aws/sagemaker/Endpoints | Host vCPU and RAM use per instance |
GPUUtilization, GPUMemoryUtilization | /aws/sagemaker/Endpoints | GPU compute and GPU memory use; the second decides whether a smaller GPU still fits |
Read the metrics as a budget. If users wait seconds but ModelLatency plus OverheadLatency is only tens of milliseconds, the missing time is outside the endpoint: an API Gateway integration, a Lambda cold start or another downstream call. Scaling or profiling the endpoint will not help. Turn on X-Ray active tracing on the API Gateway stage and the Lambda function; each trace then shows a segment per hop with its duration. Also note that MemoryUtilization is host RAM: before moving a model to a GPU with less memory, check GPUMemoryUtilization, because low GPUUtilization says nothing about whether the weights fit.
On a serverless endpoint, slow first requests after a quiet period that coincide with ModelSetupTime spikes (while ModelLatency is steady) are cold starts. Provisioned concurrency keeps environments initialised and is billed while configured; raising memory size or maximum concurrency does not remove setup time.
X-Ray sampling is decided at the entry point
X-Ray does not trace every request: sampling rules (the default traces the first request each second plus 5% of the rest) decide which requests are recorded, and the decision is made by the first instrumented service a request reaches and passed downstream in the trace header. A rule that matches only a downstream service therefore never applies when that service is always called by an instrumented upstream. To trace more of one path, create a higher-priority rule that matches the entry service and the specific URL path, rather than raising the default rule (which samples everything more) or tuning a rule on the downstream service.
Troubleshooting agents with AgentCore Observability
Amazon Bedrock AgentCore Observability shows, for each agent session, a trace whose spans record every model invocation, tool call and step with its duration and order, alongside runtime metrics, in the CloudWatch generative AI observability console. That per-session timeline is what you need when some sessions are slow: aggregate metrics blend all sessions together, and a log of model calls alone shows only one kind of step, so neither reconstructs the sequence of steps that made one session slow.
Agents hosted on AgentCore Runtime publish metrics automatically. Spans need two things:
- Instrumentation in the agent: the AWS Distro for OpenTelemetry (ADOT) SDK in its dependencies, with the agent started through
opentelemetry-instrument(frameworks such as Strands emit spans through it). The separate ADOT Collector is not supported for AgentCore agents. - CloudWatch Transaction Search enabled once per account, so trace segments are ingested into CloudWatch Logs as spans. If metrics appear but no traces appear for any session, this is the missing step; sampling rules change how many requests are traced, not whether any are stored.
The monitoring of answer quality, online evaluations and the GenAI observability prebuilt views as quality tools belong to Task 4.1; here the focus is finding where time and resources go.
How do you build performance dashboards in CloudWatch?
A useful performance dashboard plots the right statistic, covers every account and resource automatically, and does not copy data. Three CloudWatch features do most of the work.
- Percentile statistics. Latency complaints are about the slowest requests.
Averageover thousands of fast requests stays flat when a few take seconds; plotp90,p95orp99to show the tail.SumandSampleCountmeasure volume, and a longer period smooths spikes away. - Cross-account observability. Link workload (source) accounts to a central monitoring account with CloudWatch cross-account observability (Organizations or individual links). Dashboards, alarms and queries in the monitoring account can then use the linked accounts'
AWS/SageMaker,AWS/Bedrockand other metrics, logs and traces directly, with no metric streams, log subscriptions or extra services. Per-account dashboards with links between them are not a single view. - Metric math and
SEARCH. A widget built from fixed metric IDs only shows the resources that existed when it was made. ASEARCHexpression, for exampleSEARCH('{AWS/SageMaker,EndpointName,VariantName} MetricName="ModelLatency"', 'Average', 60), re-runs when the dashboard loads, so new endpoints that match appear without anyone editing it. Other metric math (such asSUM, ratios orMETRICS()) derives values like error rate or tokens per invocation.
Useful generative AI metrics in AWS/Bedrock include Invocations, InvocationLatency, InvocationThrottles, InputTokenCount, OutputTokenCount, CacheReadInputTokenCount and CacheWriteInputTokenCount, by ModelId. For alerting without fixed thresholds, a CloudWatch alarm based on anomaly detection learns a metric's hourly and daily pattern and fires within minutes when, for example, OutputTokenCount leaves its expected band.
How do you optimise endpoint capacity for cost, performance and reliability?
Capacity optimisation means running just enough instances, in the right shape, at the right time, while meeting the availability requirement. These levers recur:
| Requirement | Lever | Trap |
|---|---|---|
| Survive the loss of one Availability Zone | Run two or more instances in the production variant; SageMaker AI spreads them across AZs. To keep cost flat, split one large instance into two smaller ones that each carry the full load (an instance running well under half its capacity can become two half-size instances, each still able to carry the whole load alone) | Auto scaling from one instance only replaces capacity after a failure; a second Region adds cost and complexity the requirement does not need |
| Predictable daily peak with launch lag | Application Auto Scaling scheduled actions that raise MinCapacity shortly before the peak and lower it afterwards, alongside target tracking | Lowering the target value or step scaling still reacts after traffic arrives; a high floor all day wastes the overnight hours |
| Many small, rarely called models with the same framework | Multi-model endpoint: one shared fleet loads each model from Amazon S3 on first use and evicts idle ones | Multi-container endpoints host up to 15 distinct containers, not hundreds of models; a Savings Plan only discounts idle capacity |
| Long-running requests, no caller waiting | Asynchronous inference, which can scale to zero instances | Not suitable when callers need a synchronous response |
| Intermittent traffic, CPU model | Serverless inference (pay per use; provisioned concurrency for cold starts) | Serverless inference does not support GPUs |
The trade-off is always explicit: a multi-model endpoint accepts a load delay when an evicted model is called; scale-to-zero accepts a cold start; a scheduled floor accepts paying for instances slightly before they are busy. Choose the lever whose penalty the scenario says is acceptable. Target-tracking metrics and scaling policy design are taught in the Task 3.2 lesson on endpoint scaling.
Which cost management tools track and cap ML and AI spend?
AWS cost tools split into attribution (who spent it), forecasting and limits (will we overspend, and what happens then) and anomaly alerts (is something unusual). Match the tool to the job.
| Need | Tool | Key behaviour |
|---|---|---|
| Group spend by team, project or cost centre | Cost allocation tags + Cost Explorer | Tag keys appear in Cost Explorer only after they are activated as user-defined cost allocation tags in the Billing and Cost Management console (the management account in an organisation); data is available from activation onward unless you request a backfill (up to 12 months of earlier usage) |
| Attribute on-demand Amazon Bedrock usage to teams | Application inference profiles | Create one profile per team from a model (or a cross-Region profile), tag it, and invoke the profile's ARN; usage is billed against the profile's tags |
| Warn before a monthly limit is reached | AWS Budgets with a forecasted alert | Alerts when projected month-end cost exceeds the budget; actual-cost alerts and billing alarms fire only after money is spent |
| Act automatically at a threshold | AWS Budgets actions | Apply an IAM policy, apply an SCP, or stop specific EC2 or RDS instances; an IAM policy that denies sagemaker:CreateTrainingJob and sagemaker:CreateEndpoint blocks new work without stopping running jobs |
| Detect unusual spend without thresholds | AWS Cost Anomaly Detection | Machine-learning monitors over billing data, with individual or daily/weekly summary alerts; billing data lags, so it suits finance, not paging |
| Detect unusual usage within minutes | CloudWatch anomaly detection alarm on a usage metric (for example OutputTokenCount) | Near real time, learns the pattern, no fixed threshold |
| Hard ceilings on consumption | Service quotas (instance counts, Bedrock tokens and requests per minute) | Limits throughput per account and Region; they cap capacity, not spend |
Two attribution details matter for Bedrock. Bedrock also supports cost allocation by IAM principal tags, which separates teams when each team calls through its own role or user; when teams all call through one shared role, that role's tags cannot tell them apart, and per-team application inference profiles are the dependable split. And tags only become cost dimensions once activated: tagging resources alone never makes a tag appear in Cost Explorer. Budgets actions cannot stop SageMaker AI resources directly, so a deny policy on the create actions is the way to freeze new spend while leaving running work alone.
Which purchasing options lower ML infrastructure cost?
Commitment discounts pay off for steady usage, and the right one depends on which service the usage runs on. For SageMaker AI the commitment product is the SageMaker Savings Plan.
| Option | Covers | Flexibility |
|---|---|---|
| SageMaker Savings Plans | Eligible SageMaker AI instance usage, including training, real-time inference, batch transform, processing and notebook/Studio instances | Any instance family, size, Region and component; 1- or 3-year term with an hourly spend commitment |
| Compute Savings Plans | EC2, Fargate and Lambda | Across families and Regions, but not SageMaker AI |
| EC2 Instance Savings Plans / Reserved Instances | EC2 in one family (and Region) | Least flexible; do not apply to SageMaker AI |
| Managed Spot Training | SageMaker AI training jobs on spare capacity | Large discount; jobs can be interrupted, so use checkpoints (taught in the Task 2.2 training lesson) |
| On-demand | Everything | No commitment; right for spiky, short-lived or uncertain usage |
Size the commitment to the baseline. A Savings Plan commitment is applied hour by hour, and any commitment not used in an hour is lost. If usage never drops below a steady baseline but rises for a few hours a day, commit to cover the baseline and run the peak at on-demand rates. Committing to the daily average or the peak looks like more discount, but the unused commitment in every off-peak hour costs more than the extra discount earns. A commitment only discounts capacity; it never fixes capacity that should not exist, such as hundreds of idle endpoints or an oversized GPU instance.
What drives the cost of foundation model inference in production?
The cost of FM inference depends on how you buy capacity: per token, per tier, per batch, per reserved unit or per instance-hour. Each option has a different break-even point.
| Option | How you pay | Fits |
|---|---|---|
| Bedrock on-demand (Standard tier) | Per input and output token | Variable or low traffic; no idle cost |
Bedrock Flex tier (service_tier = flex) | Discounted per-token price; requests may be processed more slowly | Latency-tolerant synchronous work, such as overnight agentic jobs, with no change to how calls are made |
| Bedrock Priority tier | Premium per-token price | Latency-critical traffic that should be served first |
| Bedrock Reserved tier and Provisioned Throughput | Committed capacity billed by time, whether used or not | Steady high volume, guaranteed capacity, and (Provisioned Throughput) many custom models |
| Bedrock batch inference | Discounted per-token price; JSONL input and output in Amazon S3 | Large offline jobs where no caller waits |
| Self-hosting on a SageMaker AI endpoint | Per instance-hour, regardless of traffic | Sustained high utilisation, custom or open-weight models not offered serverlessly, or special runtime needs |
The service tier is set per request, and not every model supports every tier. Two scenario patterns follow. When jobs must remain ordinary synchronous calls without reworking input into S3 files, Flex is the discount that fits; batch inference is cheaper only if the work can move to S3 jobs. When a self-hosted model sits on a GPU instance at a few percent utilisation, most of the hourly charge is waste; if the same model is offered as a Bedrock on-demand model in the required Region, per-token billing removes the idle cost, whereas a Savings Plan would only discount the waste and SageMaker serverless inference is not an option for GPU models.
Cross-Region inference changes where requests run, not just capacity. A geographic inference profile routes requests to other Regions in the same geography (priced as the source Region), and global profiles can route anywhere; neither satisfies a rule that requests must stay in one Region. Model selection, prompt routing and the on-demand versus Provisioned Throughput deployment decision are covered in the Task 2.1 and 3.1 lessons.
How do you monitor and reduce agent resource consumption?
An agent costs two separate things: model tokens (billed by Bedrock per model call) and runtime compute (billed by AgentCore Runtime for the session's vCPU and memory). Monitor them separately, because a looping agent inflates tokens while an idle session inflates memory.
AgentCore Runtime publishes CloudWatch metrics including Invocations, Session Count, Latency, CPUUsed-vCPUHours and MemoryUsed-GBHours, with dimensions for the service, the agent resource and the endpoint name, plus per-session usage logs. To answer which agent endpoints drive the Runtime bill, group the vCPU-hour and GB-hour metrics by endpoint; invocation counts, session counts or latency multiplied by calls are proxies that ignore how Runtime bills.
How Runtime bills: CPU is charged only for active processing (time spent waiting on a model response or other I/O is not CPU time), while memory is charged for as long as the session exists, including idle periods, until it terminates. A session ends when the client calls StopRuntimeSession, when it has been idle for idleRuntimeSessionTimeout (default 900 seconds) or when it reaches maxLifetime (default 28,800 seconds).
- Sessions that last far longer than conversations, with low vCPU use, mean memory is the cost: call
StopRuntimeSessionwhen the user ends the chat, and loweridleRuntimeSessionTimeoutfor clients that disappear without calling it. - The session lifetime settings and explicit stops control memory cost; token limits control model cost, not session memory.
- Reuse one
runtimeSessionIdfor every turn of a conversation: each new ID starts a new session, which loses the conversation's context and multiplies session start-ups.
For the token side, watch OutputTokenCount and Invocations per model, cap agent iterations in the framework, and alert on anomalies, as described above.
How do you control token usage and quota burn?
Token cost and throttling are both driven by how many tokens each request sends, asks for and repeats. Four levers cover most cases.
- Right-size the maximum output tokens. When a request starts, Bedrock deducts the input tokens plus the requested maximum output tokens from the tokens-per-minute quota, and only corrects to the actual output when the response finishes. A service whose maximum is set far above what it typically writes reserves many times what it uses and is throttled while CloudWatch shows usage far below the quota. Because the reservation is set by the requested maximum, setting that maximum just above the longest expected response is what fixes it without a quota increase. For some models, such as recent Anthropic Claude models, each output token also counts against the quota at a higher burndown rate (5x) than an input token, which makes oversized outputs burn quota even faster.
- Prompt caching for repeated prefixes. Put a cache checkpoint (
cachePointin the Converse API) after the content that is identical on every request, such as system instructions and tool definitions, and before the part that changes. Cache reads are billed at a reduced rate. A checkpoint placed after the per-request message produces a cache write on every request and no reads (CacheWriteInputTokenCountrising,CacheReadInputTokenCountat zero). Each model has a minimum number of tokens per checkpoint (for example 1,024 for Claude Sonnet 4.5); a shorter prefix is silently not cached and the cache counters stay at zero. Caching works with Converse, ConverseStream and InvokeModel on on-demand inference, and cache entries expire after a short idle period. - Trim the conversation history. Resending the full history makes
InputTokenCountgrow with every turn. Keep recent turns verbatim and replace older ones with a running summary. - Retrieve fewer, better passages. Each retrieved passage adds input tokens on every turn. Lower
numberOfResultson the knowledge baseRetrievecall after testing that answers stay good.
Input tokens fall only when less text is sent, so history length and retrieval size are the input-side levers. Choosing a cheaper model or a prompt router for simple requests is the other big lever (Task 2.1).
How do you control embedding and vector storage costs?
RAG systems add two AI-specific cost lines: computing embeddings (tokens through an embedding model) and storing and searching vectors. Both scale with how much text you embed and how big each vector is.
Embedding computation
- Sync incrementally. A Bedrock knowledge base ingestion job on an existing data source detects added, modified and deleted documents and embeds only those. Deleting and re-creating the data source forces every document to be parsed, chunked and embedded again.
- Use batch inference for bulk one-off embedding. Amazon Titan Text Embeddings V2 supports Bedrock batch inference: JSONL records in S3, embeddings written back to S3, at the discounted batch rate. The per-token price is set by how Bedrock is invoked, not by where the calling code runs.
- Embed less text. Chunk overlap repeats tokens: the higher the overlap, the more text is embedded more than once, so keep it modest. S3 inclusion prefixes keep superseded or irrelevant objects out of ingestion. Semantic chunking calls a foundation model to find boundaries, which adds cost; hierarchical chunking changes what retrieval returns rather than how much is embedded. Requesting fewer output dimensions does not lower the input-token price.
Vector storage
| Store | Cost model | Fits |
|---|---|---|
| Amazon S3 Vectors | Pay for storage and queries; no clusters or capacity units | Very large vector sets with low query volume where sub-second latency is acceptable |
| Amazon OpenSearch Serverless | OpenSearch Compute Units (OCUs) for indexing and search, plus storage | Low-latency, high-throughput search and hybrid search |
| Aurora PostgreSQL with pgvector | A running database cluster | Vectors alongside relational data |
OpenSearch Serverless capacity is set per collection group: collections in one group share its OCUs, each group has minimum and maximum OCUs for indexing and search, the minimum can be 0 (accepting slower first requests after idle periods), and a group holds one collection type. Pooling idle vector collections into one group and setting minimums to 0 cuts the bill without capping daytime capacity; lowering the maximum would cap it.
To shrink the index itself, shrink each vector. Titan Text Embeddings V2 outputs 1,024, 512 or 256 dimensions, and can return binary embeddings (1 bit per dimension instead of a 32-bit float) that OpenSearch Serverless stores in a binary vector index. Both need re-embedding and a matching index, and both trade a little retrieval precision.
Tip. Task 4.2 questions describe a model, foundation model, RAG pipeline or agent in production with a cost, latency or capacity symptom, usually backed by metric values, and ask which configuration, metric, tool or purchasing option fixes it. The deciding detail is often a mid-scenario constraint: no custom benchmarking harness, no copying of metrics, running work must not stop, requests must stay in one Region, calls must remain synchronous without S3 input files, keep cost close to today's, no fixed thresholds, or newly created resources must appear automatically. Expect near-twin options that differ in one component: ModelLatency versus OverheadLatency, MemoryUtilization versus GPUMemoryUtilization, default versus advanced Inference Recommender jobs, SageMaker versus Compute Savings Plans, baseline versus average commitment, Flex versus Priority versus batch, multi-model versus multi-container endpoints, Budgets forecast versus actual alerts, Cost Anomaly Detection versus CloudWatch anomaly detection, idle timeout versus maximum lifetime, and cache checkpoint placement. Some questions combine two facts, such as GPU memory per instance size, X-Ray sampling at the entry service, or how a commitment applies hour by hour.
- Pick the instance family that supplies the bottleneck resource: c for CPU-bound, r for memory-heavy, GPU only when the model uses it, and check GPU memory (16, 24 or 48 GB) before LLM fit.
- Inference Recommender default jobs shortlist instances; advanced load test jobs take your traffic phases, endpoint configurations and percentile-latency stopping conditions.
- ModelLatency plus OverheadLatency is the time inside SageMaker AI; the rest of a slow request is found with X-Ray active tracing, whose sampling is decided at the entry service.
- AgentCore traces need the ADOT SDK and CloudWatch Transaction Search enabled once per account; metrics appear without it, spans do not.
- Dashboards: plot percentiles for latency, use cross-account observability for one central view, and SEARCH expressions so new resources appear automatically.
- Tags reach Cost Explorer only after activation; per-team Bedrock attribution uses tagged application inference profiles; Budgets forecast alerts warn early and Budgets actions apply IAM policies or SCPs.
- SageMaker AI commitments are SageMaker Savings Plans (not Compute Savings Plans or EC2 RIs), sized to the steady hourly baseline.
- Bedrock Flex discounts latency-tolerant synchronous calls; batch discounts S3 jobs; Priority costs more; idle GPU endpoints are better replaced by per-token on-demand where the model is offered.
- AgentCore Runtime bills CPU only while processing but memory until the session ends, so stop sessions and shorten the idle timeout.
- Bedrock reserves input plus maximum output tokens against the quota; cache checkpoints go after the static prefix and must meet the model's minimum size; resync incrementally and shrink vectors to cut RAG cost.
Frequently asked questions
Do Compute Savings Plans cover Amazon SageMaker AI?
No. Compute Savings Plans cover Amazon EC2, AWS Fargate and AWS Lambda. SageMaker AI usage, including training, real-time inference, batch transform, processing and notebook instances, is discounted by SageMaker Savings Plans, which apply across instance families, sizes and Regions for a 1- or 3-year hourly commitment.
How should I size a SageMaker Savings Plan commitment?
Size it to the steady baseline that your usage never drops below. A Savings Plan commitment applies hour by hour and unused commitment in any hour is lost, so committing to the daily average or the peak wastes money in off-peak hours. Run usage above the baseline at on-demand rates.
Why is my Amazon Bedrock application throttled when token usage is below the quota?
Bedrock deducts the input tokens plus the requested maximum output tokens from the tokens-per-minute quota when each request starts, and only adjusts to actual usage when it finishes. If the maximum output tokens setting is far larger than real responses, most of the quota is reserved but unused. Lower the maximum to just above the longest expected response.
Why does Bedrock prompt caching show no cache reads?
Either the cache checkpoint is placed after content that changes on every request, so each request writes a new cache entry and never reads one, or the prefix before the checkpoint is shorter than the model's minimum tokens per checkpoint (for example 1,024 for Claude Sonnet 4.5), in which case Bedrock skips caching. Place the checkpoint after the static instructions and tool definitions.
How is Amazon Bedrock AgentCore Runtime billed?
AgentCore Runtime charges for vCPU only while the agent is actively processing, not while it waits for model responses or other I/O, and charges for memory for as long as the session exists, including idle time, until it is stopped, times out after the idle timeout (default 15 minutes) or reaches its maximum lifetime (default 8 hours). Model tokens are billed separately by Bedrock.
How do I track Amazon Bedrock costs per team?
Create an application inference profile for each team from the model, tag it with the team, invoke the profile's ARN instead of the model ID, and activate the tag key as a user-defined cost allocation tag in the Billing and Cost Management console. Cost Explorer can then group Bedrock spend by that tag from activation onward, and you can request a backfill of up to 12 months of earlier data. If each team already calls through its own IAM role or user, Bedrock cost allocation by IAM principal tags is another option; it cannot separate teams that share one role.
When should I use Amazon S3 Vectors instead of OpenSearch Serverless?
Use S3 Vectors for very large vector sets that are queried relatively rarely and can tolerate sub-second latency, because it charges for storage and queries with no capacity units to run. Use OpenSearch Serverless when you need low-latency, high-throughput or hybrid search and can pay for OpenSearch Compute Units.
What is the difference between an Inference Recommender default job and an advanced job?
A default job load-tests a model package on a set of candidate instance types and returns recommendations with latency, throughput and cost. An advanced load test job lets you specify the endpoint configurations to test, a traffic pattern with phases such as initial users and spawn rate, and stopping conditions such as a P95 latency threshold.
Source
This lesson covers the "Operating, Monitoring, and Securing ML and AI Solutions" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.
- AWS Certified Machine Learning Engineer – Associate (MLA-C02) exam guide — Amazon Web Services
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.