Provisioning ML Resources: Endpoints, Scaling, Knowledge Bases and Agents (MLA-C02)
Provisioning and configuring resources for ML and AI workloads means fitting the capacity, network, containers, scaling, retrieval and agent infrastructure behind a model to an existing architecture and its requirements. Task 3.2 of MLA-C02 covers choosing on-demand versus provisioned capacity on Amazon Bedrock and SageMaker AI, passing values between CloudFormation stacks and Step Functions, building inference and training containers, running SageMaker AI endpoints in a VPC, deploying with Boto3, the AWS CLI and the SageMaker Python SDK, picking auto scaling metrics, scaling GPU workloads, configuring Amazon Bedrock knowledge bases and retrieval pipelines, and hosting agents and their state on Amazon Bedrock AgentCore.
On this page10 sections
- When should you use on-demand capacity and when provisioned capacity?
- How do you automate provisioning across stacks and orchestration services?
- How do you build and maintain containers for ML workloads?
- How do you configure SageMaker AI endpoints inside a VPC?
- How do you deploy and update models programmatically?
- Which metric should an auto scaling policy track?
- How do you scale GPU resources for AI workloads?
- How do you set up an Amazon Bedrock knowledge base: vector store and document indexing?
- How do you implement a retrieval pipeline that meets business needs?
- How do you deploy agent infrastructure and manage agent state?
- Choose between on-demand, cross-Region, provisioned and batch capacity on Amazon Bedrock, and size SageMaker AI serverless provisioned concurrency to cut idle cost.
- Pass values between CloudFormation stacks with exports or Parameter Store dynamic references, and wait on SageMaker AI jobs and endpoints in Step Functions without Lambda.
- Build ML containers that meet the SageMaker AI inference and training contracts, extend AWS images, and keep Amazon ECR repositories scanned and clean.
- Configure SageMaker AI endpoints in private subnets with the right gateway and interface VPC endpoints, private DNS and network isolation.
- Deploy, update and wait for endpoints with Boto3, the AWS CLI and the SageMaker Python SDK, and select auto scaling metrics, including scale-to-zero policies and GPU metrics.
- Configure Amazon Bedrock knowledge bases (vector store, dimensions, parsing, chunking, metadata) and retrieval pipelines with server-side filters and structured data stores.
- Host agents on AgentCore Runtime with Gateway and Identity, and keep agent state in the right store.
When should you use on-demand capacity and when provisioned capacity?
Use on-demand capacity when traffic is unpredictable, bursty or idle for long stretches and you will not commit to a term; use provisioned capacity when traffic is steady and high, every request needs consistent latency, or the workload needs capacity that nothing else can consume. Skill 3.2.1 asks you to make this trade-off for both Amazon Bedrock and SageMaker AI, and the scenario's constraints (commitment, idle cost, cold starts, interactivity) decide it.
Amazon Bedrock
| Option | Billing and capacity | Fits | Does not fit |
|---|---|---|---|
| On-demand, single Region | Per token; shares the Region's capacity under account quotas (tokens and requests per minute) | Variable traffic, experimentation, no commitment | Peaks above the quota (throttling) |
| Cross-Region inference profile (geographic or global) | Still per token on demand; requests are routed across several Regions | Unpredictable bursts that throttle in one Region, when compliance allows processing in other Regions of the geography (or anywhere, for global) | A need for capacity reserved for one workload |
| Provisioned Throughput | Model units bought for a 1- or 6-month commitment term (custom and imported models can also be bought with no commitment, billed hourly); billed whether used or not; invoked through the provisioned model ARN | High, steady, interactive traffic; capacity no other workload can consume | Bursty or idle traffic, or a team that will not pay for idle time |
| Batch inference | Job over JSONL records in Amazon S3; output written to S3 | Large offline workloads at lower cost | Anything a user waits for |
Only Provisioned Throughput reserves model capacity for one workload, and you must send requests to the provisioned model's ARN (not the base model ID) to use it.
SageMaker AI
A real-time endpoint runs instances you choose and bills for every hour they run, idle or not. Serverless Inference bills per request and memory size, runs on CPU only (no GPUs) and can cold start after idle time. Provisioned concurrency keeps a set number of serverless workers initialised, so requests avoid the cold start, and Application Auto Scaling can manage it on a schedule or with target tracking on the predefined SageMakerVariantProvisionedConcurrencyUtilization metric. A provisioned-concurrency scalable target has a minimum of 1, not 0: to cut idle cost outside business hours, schedule it down to 1 rather than removing it. Asynchronous inference can scale its instances to zero, but callers fetch results from Amazon S3, so it cannot answer interactively. Choosing the inference type itself is Task 3.1; Task 3.2 is about sizing and paying for the capacity behind it.
How do you automate provisioning across stacks and orchestration services?
Pass values between AWS CloudFormation stacks with exported Outputs and Fn::ImportValue, or with AWS Systems Manager Parameter Store dynamic references when the producer must stay free to change the value; and drive SageMaker AI from AWS Step Functions with service integrations instead of custom polling code. Skill 3.2.2 tests which mechanism carries the value and which integration pattern waits correctly.
| Mechanism | How it works | Watch out for |
|---|---|---|
Outputs with Export + Fn::ImportValue | The producer stack publishes a named value in the account and Region; any stack imports it by name | While another stack imports an export, the producer cannot change or delete that export, so the stacks become locked together |
Parameter Store + dynamic reference ({{resolve:ssm:name}}) | The producer writes a parameter; the consumer template resolves it at deploy time | No export lock: the producer can replace the value, and consumers pick it up on their next deployment |
| Nested stack outputs | A parent template contains the child stack and reads Outputs from it | Merges ownership into one stack tree, so separately owned stacks are no longer separate |
Fn::GetAtt / Ref | Reads attributes of resources in the same template | Cannot reach into another stack |
Mappings + Fn::FindInMap | Static lookup tables written into the template | Values are copied by hand and go stale |
Step Functions integration patterns for SageMaker AI
Step Functions offers three patterns: request-response (call the API and move on), run a job (.sync, wait until the job finishes) and wait for a callback (.waitForTaskToken). The optimized SageMaker AI integration supports .sync only for jobs: training, processing, transform, hyperparameter tuning and labeling. It does not support .waitForTaskToken at all. CreateModel, CreateEndpointConfig, CreateEndpoint and UpdateEndpoint are request-response, so the state finishes as soon as SageMaker AI accepts the request, while the endpoint is still Creating.
To wait for an endpoint without AWS Lambda, build a loop: a Wait state, an AWS SDK integration task that calls DescribeEndpoint, and a Choice state that continues on InService, goes to a Fail state on Failed, and otherwise returns to the Wait. The principle: branch on the status the service reports, and treat a failed state as terminal.
How do you build and maintain containers for ML workloads?
Start from an AWS prebuilt image whenever you can, add only what it lacks, and store the result in Amazon ECR; a container that SageMaker AI runs must follow its fixed contract for ports and paths. Skill 3.2.3 covers the contract, the ways to customise images and how to keep repositories clean and scanned.
The inference container contract
SageMaker AI starts a hosting container and sends GET /ping health checks and POST /invocations inference requests to port 8080. A server on another port or with its own routes (for example /predict or /health) is never reached, so the endpoint fails its health checks and never becomes InService. The contract is all-or-nothing: the routes and the port must both match. Model artifacts from model.tar.gz are extracted to /opt/ml/model in the container.
The training container paths
| Path | Contents |
|---|---|
/opt/ml/input/data/<channel> | Training data for each input channel |
/opt/ml/input/config | Hyperparameters and resource configuration |
/opt/ml/model | Whatever the script saves here is packed into model.tar.gz and uploaded to the job's S3 output path |
/opt/ml/checkpoints | Checkpoints synced to S3 when checkpointing is configured (used with Spot training) |
A model saved to /tmp or any other path is discarded with the container, even though the job reports Completed. No AWS SDK upload code is needed.
Customising a prebuilt image
- requirements.txt beside the inference script: the prebuilt framework containers pip-install it when the container starts. It adds Python packages only (no OS packages), and it needs to reach a package repository at startup: the internet, or a private mirror such as AWS CodeArtifact through a VPC endpoint.
- Extend the AWS image: a Dockerfile
FROMthe AWS Deep Learning Container or framework inference image that installs OS and Python packages, pushed to ECR. Everything is baked in, and you keep AWS's serving stack and its patched base. - Build from scratch (for example from a plain CUDA image): works, but your team now maintains the model server and every patch.
Maintaining images in Amazon ECR
| Feature | What it does |
|---|---|
| Basic scanning | Scans operating system packages, on push or manually |
| Enhanced scanning (Amazon Inspector) | Scans OS and programming-language packages (for example Python); with continuous scanning, findings refresh when new CVEs are published for images already stored |
| Lifecycle policy | Expires images by rule, for example untagged images older than a number of days, with no scripts |
| Tag immutability | Stops a tag being overwritten; it removes nothing |
| Replication | Copies images to other Regions or accounts; it neither scans nor cleans up |
How do you configure SageMaker AI endpoints inside a VPC?
Put the subnets and security groups in the VpcConfig of the SageMaker AI model (CreateModel), not in the endpoint configuration, and then give those private subnets a private path to every AWS service the model and its callers need. Skill 3.2.4 tests which VPC endpoint carries which traffic.
The model's network interfaces live in your subnets. With no NAT gateway, the endpoint cannot even download its artifacts from Amazon S3 until the subnets' route tables have an S3 gateway endpoint. Anything the inference code calls (Amazon DynamoDB, Amazon Bedrock, others) also needs an endpoint.
| Need | Endpoint | Type |
|---|---|---|
| Model artifacts or data in Amazon S3 | com.amazonaws.<region>.s3 | Gateway (route table entry) or interface |
| Amazon DynamoDB | com.amazonaws.<region>.dynamodb | Gateway (or interface) |
Apps calling InvokeEndpoint | com.amazonaws.<region>.sagemaker.runtime | Interface |
Control-plane calls (CreateModel, DescribeEndpoint) | com.amazonaws.<region>.sagemaker.api | Interface |
Amazon Bedrock InvokeModel, Converse | com.amazonaws.<region>.bedrock-runtime | Interface |
| Amazon Bedrock control plane (model management) | com.amazonaws.<region>.bedrock | Interface |
Three rules decide most questions. First, only Amazon S3 and Amazon DynamoDB offer gateway endpoints; every other service, including SageMaker AI and Amazon Bedrock, uses an interface endpoint (AWS PrivateLink). Second, runtime and control-plane APIs are separate endpoint services: an endpoint for bedrock or sagemaker.api does not carry bedrock-runtime or sagemaker.runtime calls. Third, code that uses the AWS SDK's default hostnames reaches an interface endpoint only when private DNS is enabled on it, so those names resolve to the endpoint's private IP addresses; and the endpoint's security group must allow HTTPS (port 443) from the caller's security group.
Network isolation
EnableNetworkIsolation is a CreateModel parameter that removes all network access from the container, including calls to AWS services, whatever the code inside tries. SageMaker AI itself still pulls the image and model artifacts, so the endpoint starts normally. It is the control for untrusted or vendor code. Security groups and route changes only filter traffic inside the VPC: they can block the artifact download or leave VPC endpoints reachable, so they do not prove the container is silent. Network isolation is not a way to reach S3 privately; that is the gateway endpoint's job.
How do you deploy and update models programmatically?
Deploying with Boto3 or the AWS CLI takes three calls in order: create_model, then create_endpoint_config, then create_endpoint. Skill 3.2.5 tests the order, what each resource holds, how to swap a model without downtime and how to wait for an endpoint.
| Resource | Holds |
|---|---|
Model (CreateModel) | Container image URI, model artifact S3 location, execution role, VpcConfig, network isolation |
Endpoint configuration (CreateEndpointConfig) | Production variants: model name, instance type and count (or serverless settings), variant weights, async and data capture settings |
Endpoint (CreateEndpoint) | The named, invokable resource, created from one endpoint configuration |
create_endpoint takes an endpoint configuration name, not an instance type. Endpoint configurations are immutable: to move an endpoint to a new model version, create a new model and a new endpoint configuration, then call update_endpoint with the new configuration. SageMaker AI provisions the new fleet, shifts traffic and keeps the same endpoint name in service throughout. Deleting and recreating the endpoint causes downtime, and create_endpoint with an existing name fails.
Creation is asynchronous: create_endpoint returns while the endpoint is Creating. To block until it can serve traffic, use the SageMaker AI client's endpoint_in_service waiter with the endpoint name (CLI: aws sagemaker wait endpoint-in-service --endpoint-name ...). The waiter polls DescribeEndpoint and stops on InService or errors on Failed. The waiter needs the name of the resource whose status it reads, which is the endpoint.
The SageMaker Python SDK wraps the same three calls: model.deploy(instance_type=..., initial_instance_count=...) creates the model, configuration and endpoint and returns a predictor, waiting for the endpoint by default. Infrastructure-as-code tools create the same three resources.
Which metric should an auto scaling policy track?
Track the metric that rises and falls in proportion to the load each instance or copy carries: invocations per instance for short, uniform requests; concurrent requests for long streaming requests; backlog per instance for asynchronous endpoints; and a resource metric such as GPU utilisation when requests differ widely in cost. Skill 3.2.6 tests this choice and the policies that make scale-from-zero work.
SageMaker AI scaling runs through Application Auto Scaling. The scalable dimension is the variant's instance count, an inference component's copy count, or a serverless variant's provisioned concurrency. Target tracking accepts predefined metrics directly; anything else needs a customized metric specification.
| Metric | Kind | Use it when |
|---|---|---|
SageMakerVariantInvocationsPerInstance | Predefined | Short requests whose count tracks load (typical CPU models) |
SageMakerVariantConcurrentRequestsPerModelHighResolution (and SageMakerInferenceComponentConcurrentRequestsPerCopyHighResolution) | Predefined, emitted every 10 seconds | Long or streaming LLM requests: it counts in-flight requests until the last token is sent, so scale-out reacts faster |
SageMakerVariantProvisionedConcurrencyUtilization | Predefined | Serverless endpoints with provisioned concurrency only |
ApproximateBacklogSizePerInstance | Customized | Asynchronous endpoints: queue depth per instance |
CPUUtilization, GPUUtilization, GPUMemoryUtilization | Customized | Requests vary so much in cost that counts mislead; GPU utilisation shows how busy the GPU is, while GPU memory mostly reflects the loaded model |
Mind the CloudWatch namespace in a customized metric. Invocation metrics (Invocations, ModelLatency, concurrency) are in AWS/SageMaker; instance metrics (CPUUtilization, MemoryUtilization, GPUUtilization, GPUMemoryUtilization) are in /aws/sagemaker/Endpoints. A specification with the right metric name in the wrong namespace finds no data. Endpoint-wide totals such as Invocations do not change with capacity, so they make poor target tracking metrics, and latency is an outcome rather than a load signal.
Scaling to and from zero
Target tracking cannot scale out from zero, because with no capacity there is no load metric to track. Each mechanism has its own wake-up signal:
- Inference components (real-time): set
MinInstanceCountto 0 in managed instance scaling, give each component a scalable target with minimum 0 and target tracking, and add a step scaling policy driven by a CloudWatch alarm onNoCapacityInvocationFailures, which adds a copy when a request finds no capacity. Requests that arrive while the first copy starts are rejected. Real-time scale-to-zero is available only for endpoints that host inference components. - Asynchronous endpoints: scalable target minimum 0, target tracking on backlog per instance, plus a step scaling policy on
HasBacklogWithoutCapacity; queued requests wait rather than fail. - Serverless endpoints scale to zero by design (no GPUs); provisioned concurrency floors at 1.
How do you scale GPU resources for AI workloads?
Scale GPU inference with inference components that declare each model's accelerator and memory needs plus managed instance scaling on the variant, and secure scarce GPU capacity for training with SageMaker AI training plans. Skill 3.2.10 tests both halves: packing and scaling models on GPU instances, and getting the GPUs at all.
Inference components
An inference component deploys one model onto an endpoint with its own compute requirements: number of accelerator devices, minimum memory and CPU cores. Each component scales its copy count independently with its own traffic, and SageMaker AI places copies on instances with free GPUs. Managed instance scaling on the production variant (minimum and maximum instance counts) adds instances when new copies do not fit and removes instances that are no longer needed. This is how several fine-tuned models share multi-GPU instances, each with dedicated GPUs.
| Approach | GPU behaviour |
|---|---|
| Inference components + managed instance scaling | Per-model accelerators and memory; independent copy scaling; instances follow placement needs |
| Multiple production variants | Split traffic by weight between variants (A/B tests); each variant scales instances, not per-model GPUs |
| Multi-model endpoint | Loads artifacts on first request and evicts them from shared memory; nothing reserves a GPU per model |
| Fixed peak instance count | Pays for peak all the time and cannot react |
For GPU workloads whose request cost varies, scale on GPUUtilization or concurrency rather than raw invocations (see the metrics section).
Getting GPU capacity
Read the error before choosing a fix. A service quota error means the account is not allowed more instances; an insufficient capacity error means AWS has none to give right now, and a quota increase alone does nothing. A SageMaker AI training plan reserves a chosen accelerated instance type and count for a chosen time window, for one target resource chosen at purchase: training jobs, SageMaker HyperPod clusters, SageMaker AI inference endpoints or SageMaker Studio apps. The target cannot be switched later, and the account still needs service quotas that cover the plan's instance count. Managed Spot training is cheaper but can be interrupted, and a SageMaker Savings Plan lowers the price of usage without reserving any capacity. Purchasing options and cost tooling in general belong to Task 4.2.
How do you set up an Amazon Bedrock knowledge base: vector store and document indexing?
An Amazon Bedrock knowledge base turns documents into searchable vectors: it parses each source file, splits it into chunks, embeds each chunk with an embedding model and writes the vectors with their text and metadata to a vector store you choose. Skill 3.2.7 tests the vector store choice, the index settings that must match the embedding model, and the parsing, chunking and metadata settings that control what is indexed.
Choosing the vector store
| Vector store | Fits | Notes |
|---|---|---|
| Amazon S3 Vectors | Large corpora queried infrequently; lowest storage cost; nothing to provision | Sub-second queries; no hybrid search |
| Amazon OpenSearch Serverless | Fully managed vector search with no clusters to size; frequent queries | Supports hybrid (keyword + semantic) search; Bedrock creates the index with the faiss engine |
| Amazon OpenSearch Service managed cluster | Teams already running OpenSearch domains | You size and manage the cluster |
| Amazon Aurora PostgreSQL (pgvector) | Teams that want vectors next to relational data | Supports hybrid search when the table has a text-searchable chunks column; Bedrock connects through the RDS Data API with a Secrets Manager secret |
| Amazon Neptune Analytics | GraphRAG: retrieval that follows entity relationships | Graph plus vector search |
| Third-party (Pinecone, MongoDB Atlas, Redis Enterprise Cloud) | Existing investments | MongoDB supports hybrid search with a filterable text field |
Hybrid search combines keyword and semantic matching, which matters when users search for exact identifiers (part numbers, docket numbers, error codes) that embeddings represent poorly. Bedrock supports it only on Amazon RDS/Aurora, Amazon OpenSearch Serverless and MongoDB with a filterable text field; on other stores a HYBRID query falls back to semantic search.
Aurora prerequisites: Amazon Bedrock does not open a database connection into your VPC. It calls the RDS Data API, which must be enabled on the cluster, and authenticates with an AWS Secrets Manager secret holding the database user's credentials, whose ARN you give the knowledge base. The table needs columns for the ID, the embedding, the chunk text and metadata, with a vector index (for example HNSW). The Data API endpoint plus the secret is the whole access path, so the cluster can stay private.
Dimensions must match. The vector field's dimension is fixed when the index is created and must equal the embedding model's output. Amazon Titan Text Embeddings V2 outputs 1,024 (default), 512 or 256 dimensions; 1,536 belongs to older models such as Titan Embeddings G1 - Text. A mismatch makes every ingestion fail while writing vectors, and the fix is a new index with the right dimension; settings that leave the dimension unchanged cannot fix it.
Parsing and chunking
Parsing and chunking are chosen when a data source is created and cannot be changed afterwards, so a change means a new data source. The default parser extracts text; for PDFs with tables, charts and diagrams, choose a foundation model parser (or Amazon Bedrock Data Automation) so those elements are extracted.
| Chunking strategy | Behaviour | Fits |
|---|---|---|
| Fixed-size | Chunks of a maximum token count with an overlap percentage | General text; the default |
| Hierarchical | Small child chunks are matched; their larger parent chunks are returned | Precise matches that need surrounding context |
| Semantic | Boundaries follow shifts in meaning between sentences | Prose with uneven topic lengths |
| No chunking | Each file is one chunk | Documents already split by the author (one FAQ entry or record per file); each file must fit the embedding model's input limit |
Metadata and sync
Custom attributes for filtering come from a sidecar file in the same S3 prefix named after the source file plus .metadata.json (for example policy.pdf.metadata.json, with a metadataAttributes object). S3 object tags, user-defined object metadata and differently named files are not read. A sync (StartIngestionJob) is incremental: it indexes new and changed files (including changed metadata) and removes deleted ones. Scheduling and automating refreshes belongs to Task 3.3.
How do you implement a retrieval pipeline that meets business needs?
Pick the knowledge base API that matches how much of the pipeline you want Amazon Bedrock to run, apply metadata filters on the server for anything security-relevant, and use a structured data store when answers need live aggregations. Skill 3.2.8 tests these choices.
| API | Returns | Use when |
|---|---|---|
Retrieve | Chunks with source location, metadata and relevance scores; no generation | You generate with your own prompt or a model outside Bedrock (for example a SageMaker AI endpoint), or apply your own score threshold |
RetrieveAndGenerate | A generated answer with citations from a Bedrock model | A managed end-to-end RAG answer is enough |
RetrieveAndGenerateStream | The same, streamed | Interactive UIs that show tokens as they arrive |
Because RetrieveAndGenerate runs generation itself, you cannot filter its retrieved chunks by your own threshold before the model sees them. Querying the vector index directly also works but rebuilds what Retrieve already manages.
Metadata filters and tenant isolation
A retrieval configuration can include metadata filters (equals, in, greaterThan, andAll, orAll and others) that limit the search to matching chunks before results are returned. For multi-tenant isolation, put an equals filter on the tenant attribute in every call, and set its value on the backend from the authenticated identity (for example the tenant claim in the user's token), never from the question text. Implicit filtering asks a model to derive filters from the user's query, which is useful for relevance but is not a security control, because the user controls the query. Isolation has to be enforced in the retrieval request itself, not requested from the model.
Structured data
Vector retrieval returns similar passages; it cannot sum, group or join, and nightly exports go stale. A knowledge base connected to a structured data store (Amazon Redshift, including data reached through the AWS Glue Data Catalog) converts a natural-language question into SQL, runs it against the live warehouse and returns the result, with no custom text-to-SQL code. Other retrieval strategies, such as reranking and choosing retrieval settings for answer quality, are part of Task 3.1.
How do you deploy agent infrastructure and manage agent state?
New agents are hosted on Amazon Bedrock AgentCore: AgentCore Runtime runs the agent container, AgentCore Gateway publishes tools over the Model Context Protocol (MCP), AgentCore Identity handles credentials, and AgentCore Memory keeps state that outlives a session. Skills 3.2.9 and 3.2.11 test how these parts are configured and where state must live. Amazon Bedrock Agents is now Bedrock Agents Classic and closed to new customers (from 30 July 2026), so it is a choice only for existing agents; Knowledge Bases and Guardrails are unaffected.
AgentCore Runtime
- The agent ships as a container image in Amazon ECR built for linux/arm64 (for example with
docker buildx); an x86 image fails to start. The container serves/invocationsand/pingon port 8080. - Each session runs in its own isolated microVM. A session ends after its idle timeout (default 15 minutes) or maximum lifetime (up to 8 hours on microVMs), set by
idleRuntimeSessionTimeoutandmaxLifetime. A session resumed with the sameruntimeSessionIdafter that starts on a new microVM, so in-process variables and local files are gone. - Network mode: PUBLIC reaches the internet but not your VPC. VPC mode creates network interfaces in your subnets with private IP addresses only, so attaching them to public subnets gives no internet access; outbound internet needs private subnets with a default route to a NAT gateway. An AgentCore interface VPC endpoint carries calls into the AgentCore APIs, not the agent's outbound calls.
AgentCore Gateway and Identity
Gateway turns AWS Lambda functions, OpenAPI or Smithy APIs into MCP tools behind one endpoint that many agents can share, so tools are not copied into each agent. Identity covers two directions. Inbound authorization decides who may invoke the agent (IAM or a JWT from an identity provider such as Amazon Cognito). Outbound authorization lets the agent call other services: OAuth2 credential providers (for example Google) run the user consent flow and keep each user's tokens in a token vault, so the agent gets a fresh access token on later sessions without custom storage or refresh code. Building that yourself means writing token storage, refresh and consent handling.
Where state lives
| State store | Holds | Lifetime |
|---|---|---|
| Runtime session (process memory, microVM files) | Working context within one session | Until idle timeout or maximum lifetime |
| AgentCore Memory short-term (events) | The exact conversation turns, stored as events under an actor ID and session ID | A configurable event expiry; list events to replay a conversation |
| AgentCore Memory long-term (strategies) | Insights extracted from events by strategies such as semantic facts, user preferences and summaries, scoped by actor ID | Persistent across sessions; retrieved as memory records |
Agents Classic sessionAttributes | Key-value pairs that persist across every InvokeAgent call in a session and reach the action group Lambda | The session |
Agents Classic promptSessionAttributes | Key-value pairs for one turn's prompt | One turn |
Choose short-term events when the agent must replay exactly what was said; choose a long-term strategy when it must remember distilled facts such as preferences in new sessions. A Memory resource with no strategy configured produces no long-term records. Deployment pipelines for agents belong to Task 3.3.
Tip. Task 3.2 questions describe an existing architecture (private networking, separately owned infrastructure stacks, a custom container, an endpoint that scales poorly, a knowledge base returning wrong chunks, an agent that loses state) and ask which configuration fixes it. One stated constraint usually decides between options that would all work, such as no outbound internet, no custom code, no idle cost or no long-term commitment, so read the constraints before the options. Expect options that look alike and differ in a single detail, such as the API, endpoint type, resource name or metric they use; knowing exactly what each service, endpoint and metric does is what separates them. Multiple-response items often combine two settings that each fix a separate symptom.
- Only Bedrock Provisioned Throughput reserves model capacity for one workload; cross-Region inference profiles absorb bursts while staying on-demand, and batch inference is never interactive.
- Serverless provisioned concurrency removes cold starts and can be scheduled with Application Auto Scaling, but its scalable target floors at 1, and serverless has no GPUs.
- CloudFormation exports lock the producer while imported; Parameter Store dynamic references do not. In Step Functions, SageMaker AI .sync exists only for jobs, never .waitForTaskToken, so wait for endpoints with a Wait + DescribeEndpoint + Choice loop.
- Inference containers answer GET /ping and POST /invocations on port 8080; training scripts must save the model to /opt/ml/model. Extend the AWS image FROM it when you need OS packages or have no internet.
- VpcConfig belongs on the model. Only S3 and DynamoDB have gateway endpoints; runtime and control-plane APIs (sagemaker.runtime vs sagemaker.api, bedrock-runtime vs bedrock) are separate interface endpoints, and default SDK hostnames need private DNS.
- Deploy with create_model, create_endpoint_config, create_endpoint; configurations are immutable, so swap versions with a new configuration and update_endpoint, and wait with the endpoint_in_service waiter on the endpoint name.
- Scale short requests on invocations per instance, streaming LLMs on high-resolution concurrent requests, async on backlog, and variable-cost GPU work on GPUUtilization from /aws/sagemaker/Endpoints; scale from zero needs a step policy on NoCapacityInvocationFailures (or HasBacklogWithoutCapacity for async).
- Inference components plus managed instance scaling give each model its own GPUs and scaling; training plans reserve GPU capacity, which quotas, Spot and Savings Plans do not.
- Knowledge base vector fields must match the embedding model's dimensions; hybrid search needs OpenSearch Serverless, RDS/Aurora or MongoDB; Aurora needs the Data API and a Secrets Manager secret; metadata comes from .metadata.json sidecars.
- AgentCore Runtime sessions are ephemeral ARM64 microVMs; replay exact turns from Memory events, keep distilled preferences with long-term strategies, and use Gateway for shared MCP tools and Identity for outbound OAuth tokens.
Frequently asked questions
What is the difference between Amazon Bedrock Provisioned Throughput and cross-Region inference?
Amazon Bedrock Provisioned Throughput buys model units (for base models, with a commitment term), so the capacity is reserved for your workload and billed whether used or not; you invoke it through the provisioned model ARN. Cross-Region inference keeps on-demand per-token pricing and routes requests across several Regions in a geography (or globally) to absorb bursts and reduce throttling, but it reserves nothing. Choose Provisioned Throughput for steady, high, interactive traffic that needs dedicated capacity, and cross-Region inference for unpredictable bursts with no commitment.
Why does a SageMaker AI endpoint with a custom container never reach InService?
A SageMaker AI hosting container must answer GET /ping health checks and POST /invocations requests on port 8080. If the model server listens on another port or uses its own routes, such as /predict or /health, SageMaker AI cannot reach it, the health checks fail and the endpoint never becomes InService. Fix both the port and the routes, or extend an AWS prebuilt inference image that already implements the contract.
Which VPC endpoints does a SageMaker AI endpoint in private subnets need?
A SageMaker AI model in private subnets with no NAT gateway needs an Amazon S3 gateway endpoint to download its model artifacts, plus an endpoint for every service its inference code calls, such as a DynamoDB gateway endpoint or a bedrock-runtime interface endpoint with private DNS. Applications in the VPC that call InvokeEndpoint need a sagemaker.runtime interface endpoint; the sagemaker.api endpoint carries only control-plane calls. Only S3 and DynamoDB offer gateway endpoints.
How do you update a SageMaker AI endpoint to a new model without downtime?
To update a SageMaker AI endpoint without downtime, create a model for the new version and a new endpoint configuration that references it, then call UpdateEndpoint (update_endpoint in Boto3) with the new configuration. Endpoint configurations are immutable, so editing the existing one is not possible, and deleting and recreating the endpoint takes it out of service. SageMaker AI keeps the endpoint name serving traffic while it moves to the new fleet.
How can a SageMaker AI endpoint scale to zero and back?
A real-time SageMaker AI endpoint can scale to zero only when it hosts inference components. Set MinInstanceCount to 0 in managed instance scaling, give each inference component a scalable target with a minimum of 0 and target tracking, and add a step scaling policy triggered by a CloudWatch alarm on NoCapacityInvocationFailures to add a copy when a request arrives with no capacity. Asynchronous endpoints use the HasBacklogWithoutCapacity metric for the same purpose, and serverless endpoints scale to zero on their own.
Which vector store should an Amazon Bedrock knowledge base use?
For an Amazon Bedrock knowledge base, Amazon S3 Vectors gives the lowest-cost storage with nothing to provision and suits large, rarely queried corpora; Amazon OpenSearch Serverless is fully managed and supports hybrid keyword and semantic search; Aurora PostgreSQL with pgvector keeps vectors beside relational data and also supports hybrid search; Neptune Analytics enables GraphRAG. Whichever you choose, the vector field's dimension must match the embedding model's output.
How do you stop one tenant from retrieving another tenant's documents in a Bedrock knowledge base?
To isolate tenants in a shared Amazon Bedrock knowledge base, tag every document with a tenant attribute in its .metadata.json file and pass an equals filter on that attribute in every Retrieve or RetrieveAndGenerate call. Set the filter value on the backend from the user's authenticated identity, never from the question. Implicit filtering and prompt instructions depend on user-controlled text, so they are not isolation controls.
Where should an agent on Amazon Bedrock AgentCore keep state between sessions?
An agent on AgentCore Runtime runs each session in an ephemeral microVM that ends after its idle timeout or maximum lifetime, so state that must outlive a session belongs in AgentCore Memory. Store each turn as a short-term event under the user's actor ID and session ID to replay the exact conversation later, and configure a long-term strategy, such as user preference or semantic facts, to extract insights the agent can retrieve in new sessions.
Source
This lesson covers the "Deployment and Orchestration of ML and AI Workflows" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.
- AWS Certified Machine Learning Engineer – Associate (MLA-C02) exam guide — Amazon Web Services
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.