SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
Deployment and Orchestration of ML and AI Workflows

Deploying ML Models, Foundation Models and Agents on AWS (MLA-C02)

22 min readMLA-C02 · Deployment and Orchestration of ML and AI WorkflowsUpdated

Managing deployment infrastructure for ML and AI models means choosing where and how each model serves predictions: the SageMaker AI inference option (real-time, serverless, asynchronous or batch), the compute and orchestrator behind it, how several models share an endpoint, how foundation models are hosted on Amazon Bedrock or SageMaker AI, how models trained elsewhere are brought in, how agents are deployed with their tools, and how a RAG system retrieves and reranks context. Task 3.1 of MLA-C02 tests these choices through scenarios in which one stated constraint (no idle cost, a fixed instance type, no commitment, no hosted model, existing Kubernetes tooling) decides between options that would otherwise all work.

What you’ll learn
  • Choose between real-time, serverless, asynchronous and batch inference from latency, payload, processing time and traffic pattern, and configure asynchronous endpoints and batch transform jobs.
  • Select an inference compute target (Inferentia, Graviton, NVIDIA GPU or CPU) and a deployment orchestrator (SageMaker AI, EKS, ECS or Lambda) that meet the stated constraints at the lowest cost.
  • Distinguish multi-model endpoints, multi-container endpoints, serial inference pipelines and inference components, and tune their caching and scaling.
  • Size GPU memory for a large foundation model and allocate it with tensor parallelism, quantization and per-model accelerator requests.
  • Evaluate foundation model deployment options: Bedrock on-demand, cross-Region inference, Provisioned Throughput, batch inference, Bedrock Marketplace and SageMaker JumpStart.
  • Deploy models built outside AWS with Bedrock Custom Model Import or SageMaker AI containers, and deploy agents with action groups, AgentCore Runtime, MCP and A2A.
  • Configure RAG retrieval with metadata filters, query decomposition, result counts and reranking.

Which SageMaker AI inference option fits the workload?

Choose the inference option from three facts in the scenario: whether a caller waits for the answer, how big and slow each request is, and whether traffic is steady, bursty or scheduled. Amazon SageMaker AI offers four options, and each one is the cheapest correct answer for exactly one traffic shape.

OptionHow it is calledFitsDoes not fit
Real-time endpointInvokeEndpoint; synchronous responseSteady or high traffic needing low, consistent latency; GPU modelsLong idle periods (instances bill continuously while the endpoint is in service, even when idle), large payloads, processing that takes minutes
Serverless inferenceInvokeEndpoint; synchronous responseIntermittent traffic with idle gaps where a cold start on the first call is acceptable; pay per useGPU models, strict latency on every call, large payloads; you set memory (1–6 GB) and max concurrency, not an instance type
Asynchronous inferenceInvokeEndpointAsync with an S3 InputLocation; result written to S3Large payloads (up to 1 GB) and long processing (up to about an hour), unpredictable arrivals, scale to zero instances while the queue is emptyA caller that needs the answer in the same HTTP request
Batch transformA job over an S3 prefix; output to S3A whole dataset scored on a schedule with no interactive caller; instances run only for the jobAny request-by-request or near-real-time need

Two near-twins cause most mistakes. Serverless and asynchronous inference both cost nothing while idle, but only serverless answers synchronously, and only asynchronous accepts very large payloads, long run times and GPU instances. Asynchronous inference and batch transform both read from and write to Amazon S3, but asynchronous inference handles each request as it arrives (minutes of delay), while batch transform processes a fixed dataset as one job (results when the job ends). A real-time endpoint can always be made to work, so in a cost-constrained scenario it is usually the distractor that keeps paying for idle instances.

How do you configure asynchronous inference and batch transform jobs?

Asynchronous endpoints take the payload by reference and report completion through notifications; batch transform jobs are configured by how they split, batch and filter the input data.

Asynchronous inference

  • The caller uploads the payload to Amazon S3 and calls InvokeEndpointAsync with its S3 URI as InputLocation. The call returns immediately with an OutputLocation where the result will be written. The payload is not sent in the request body.
  • To be told when each request finishes without polling S3, set success and error Amazon SNS topics in the notification settings of the endpoint's AsyncInferenceConfig. Endpoint status events describe the endpoint, not individual requests, and data capture is a monitoring feature, not a result channel.
  • An asynchronous endpoint keeps an internal queue, so it can scale to zero instances with an auto scaling policy whose minimum capacity is 0 and scale back out when requests arrive (the scaling policy itself is Task 3.2).

Batch transform data processing

  • SplitType tells the job where record boundaries are (Line for CSV or JSON Lines). Without it, each file is sent as one request, which fails with payload-too-large errors for big files.
  • BatchStrategy MultiRecord packs as many records as fit under MaxPayloadInMB into each request; SingleRecord sends one record per request. AssembleWith Line controls how outputs are joined back together, not how input is split.
  • Instances divide the work by file (S3 object), not by line. One huge file is processed by one instance however many you add; split the data into several files to parallelize.
  • InputFilter, JoinSource and OutputFilter (JSONPath expressions) run in that order: InputFilter chooses what the model receives (for example, with a record ID in the last column, $[:-2] sends every column except that ID), JoinSource Input appends the prediction to the original input record, and OutputFilter chooses what is written (for example $[-2,-1] keeps the ID and the appended prediction). Without JoinSource the output holds only predictions, so the ID cannot be kept.

Which compute target should host an inference workload?

Match the instance family to what the model actually needs: a CPU for small classical models, an inference accelerator for models that compile with the AWS Neuron SDK, and an NVIDIA GPU sized to the model's memory for everything else.

FamilyHardwareUse it whenWatch for
ml.inf2 (Inferentia2)AWS accelerator designed for inferenceHigh-volume deep learning inference (transformers, vision) at the lowest cost per inferenceThe model must compile with the Neuron SDK; custom CUDA operators rule it out
ml.trn1 / trn2 (Trainium)AWS accelerator built mainly for trainingTraining and fine-tuning; Trainium2 is also marketed for generative AI inferenceShares the Neuron SDK with Inferentia. When a scenario asks for the accelerator purpose-built for inference at the lowest cost per inference, the answer is Inferentia, not Trainium
ml.c7g / m7g (Graviton)Arm-based CPUCPU-bound models (tree ensembles, linear models) where better price-performance is the goalThe serving container must be built for arm64; an x86_64 image will not run
ml.c / m (x86 CPU)Intel or AMD CPUSmall classical models, existing x86 imagesCannot meet tight latency for large deep learning models
ml.g5 / g6 / g6eNVIDIA A10G, L4 or L40S GPUsCost-effective GPU inference; single-GPU sizes for small models, multi-GPU sizes for larger onesRight-size: pay only for the GPUs and GPU memory the model needs
ml.p4d / p5NVIDIA A100 or H100 GPUsThe largest models or the highest throughputExpensive; usually the over-provisioned distractor for small models

Right-sizing is a frequent theme: a model that needs one GPU with modest memory belongs on a single-GPU instance, not an eight-GPU one, and a small tree model gains nothing from any accelerator. Accelerators help only when the model's operators are supported by their software stack.

When should you use SageMaker AI hosting instead of EKS, ECS or Lambda?

Use SageMaker AI hosting by default when the team wants AWS to run the serving fleet; choose a container orchestrator only when an organisational or technical constraint requires it. A SageMaker AI endpoint provides a managed HTTPS API, provisions and patches instances, runs health checks, replaces failed instances and scales with an auto scaling policy, with GPU and Inferentia instance types available.

TargetChoose it whenLimits
SageMaker AI endpointLeast operational effort; ML-specific features (multi-model, inference components, async, serverless)Sits outside a Kubernetes GitOps flow
Amazon EKSThe organisation already runs everything on Kubernetes and requires the same Helm/GitOps, mesh, logging and cost tooling for models; GPU node groups host GPU modelsThe team runs the cluster and the model server
Amazon ECSThe organisation standardises on ECS; ECS on EC2 supports GPU instancesECS on AWS Fargate has no GPUs
AWS LambdaSmall, CPU-only models with light, spiky trafficNo GPUs, 15-minute execution limit, memory and package-size caps
Amazon EC2 Auto Scaling groupFull control is requiredEverything (AMIs, patching, health, scaling, load balancing) is the team's job

When a scenario states that a small team has no container or Kubernetes experience, the managed endpoint wins; when it states that every service must use the existing Kubernetes tooling, EKS wins even though SageMaker AI would be less work in isolation.

How do multi-model, multi-container, serial pipeline and inference component endpoints differ?

All four put several models behind one endpoint to share instances, but they differ in whether models share a container, whether requests pass through one model or several, and whether each model gets its own resources and scaling.

StrategyWhat it hostsHow a request is routedBest for
Multi-model endpoint (MME)Many models that share one serving container; artifacts in S3Caller sets TargetModel; the model is loaded from S3 on first use, cached, and unloaded when memory runs shortHundreds or thousands of similar, mostly infrequently used models where a cold first call is acceptable
Multi-container endpoint, Direct modeUp to 15 different containers (different frameworks) on shared instancesCaller sets TargetContainerHostname; one container per requestA few low-traffic models built with different frameworks, called independently
Serial inference pipeline2–15 containers deployed as one model (PipelineModel)Every request passes through all containers in order on the same instance; one call returns the final outputPreprocess → predict → postprocess chains
Inference componentsSeveral models (often FMs) on one endpoint, each with declared CPU, memory or accelerator requirementsCaller names the inference component; each component scales its own copy countSharing GPU instances while giving each model guaranteed resources and independent scaling, including to zero

Multi-model endpoint tuning

  • Frequent model unloads and slow responses for regularly used models mean the cache is too small: choose an instance type with more memory so more models stay loaded.
  • MMEs work best when models have similar traffic. A model with very high request rates or strict latency should move to its own dedicated endpoint so it stops evicting other models and gets consistent latency.
  • GPU multi-model endpoints are served with the NVIDIA Triton Inference Server container; each model archive (model.tar.gz) follows Triton's model-repository layout with its config.pbtxt. One archive holding every model defeats on-demand loading.

Inference component scaling

Inference components scale in two layers. Each component's copy count scales with Application Auto Scaling and can go to zero so an idle model releases its accelerators. The number of instances behind the endpoint scales with managed instance scaling on the production variant (minimum and maximum instance counts): when new copies cannot be placed because every accelerator is allocated, SageMaker AI adds instances and removes them afterwards. Production variants are a different idea: each variant runs on its own instances and TargetVariant routes traffic between them, which is for A/B testing, not for packing models together.

How do you size GPUs and allocate resources for a large foundation model?

A model fits on an instance only when its weights, plus room for the key-value (KV) cache that grows with batch size and context length, fit into the GPU memory it can actually use. Estimate weight memory as parameters × bytes per parameter: about 2 bytes in FP16/BF16, so a 7B model needs about 14 GB and a 34B model about 68 GB before any KV cache. Add headroom for the KV cache and runtime overhead before comparing with GPU memory.

  • Tensor parallelism splits each layer's weights across several GPUs on one instance. In a SageMaker AI large model inference (LMI) container, set the tensor parallel degree (for example OPTION_TENSOR_PARALLEL_DEGREE) to the number of GPUs, otherwise the model loads onto one GPU and the others' memory is wasted. Combined GPU memory (GPU count × memory per GPU) is usable only when the model is sharded across all of those GPUs; for example, four 24 GB GPUs offer 96 GB to one model only at a tensor parallel degree of 4.
  • Quantization shrinks weights when the instance is fixed: a 4-bit AWQ or GPTQ checkpoint needs roughly a quarter of the FP16 memory, at a small cost in quality. With vLLM in LMI, AWQ needs a pre-quantized checkpoint whose config the engine reads; on-the-fly options are formats such as FP8.
  • Batching settings such as the maximum rolling batch size raise throughput but also raise KV cache memory; they never make an oversized model fit. Disk (EBS) is not a substitute for GPU memory.
  • Cheapest instance that fits: compute total memory needed, find the smallest approved instance whose combined GPU memory covers it with headroom, then set tensor parallelism to use all of its GPUs.

On shared endpoints, resource allocation per model is declared on its inference component (for example the number of accelerator devices and minimum memory required), which is how several models share a multi-GPU instance without contending for the same GPU.

What are the options for deploying a foundation model on AWS?

Pick a foundation model (FM) deployment option by who manages capacity and how you pay: per token with no infrastructure, a capacity commitment, a dedicated endpoint you size, or an asynchronous batch job.

OptionWhat you manageBillingFits
Amazon Bedrock on-demand (serverless models)Nothing; call the APIPer tokenLow, unpredictable or moderate traffic; prototypes and beta products
Cross-Region inference profileNothing; invoke the profile ID instead of the model IDOn-demandThrottling at peaks with no commitment; a geographic profile (for example US or EU) keeps processing inside that geography, a global profile does not
Bedrock Provisioned ThroughputModel units purchased for a term (or no commitment for some models)Hourly for reserved capacitySteady, heavy traffic needing guaranteed throughput; required to serve some customized models
Bedrock batch inferenceJSONL prompt files in S3; outputs written to S3Discounted versus on-demand for supported modelsLarge offline jobs where nobody waits on a single response; split across jobs within per-job record quotas
Bedrock MarketplaceSubscribe, then deploy to a managed endpoint whose instance type and count you choosePer instance-hour plus any provider feeSpecialised models offered only on dedicated capacity, still called through Bedrock APIs so Bedrock tooling such as guardrails applies
SageMaker JumpStartA SageMaker AI endpoint (instance type, scaling)Per instance-hourOpen-weight models you want to control, fine-tune or host with SageMaker AI features

The common traps: Provisioned Throughput is a commitment, so it is wrong when a scenario rules commitments out; batch inference cannot serve an interactive chat; and Bedrock batch inference is not SageMaker AI batch transform, which brings back instances and containers for a model you host yourself. Data-residency rules decide between geographic and global inference profiles.

How do you deploy models that were built outside AWS?

Bring an externally trained model to AWS in one of two ways: import a supported LLM architecture into Amazon Bedrock with Custom Model Import, or host any model on SageMaker AI in a prebuilt or custom container.

Amazon Bedrock Custom Model Import

  • Supports specific open model architectures (families such as Llama, Mistral, Mixtral, Flan-T5 and Qwen; check the current list). Upload the model in Hugging Face format to Amazon S3: weights as .safetensors plus config.json and tokenizer files. A raw PyTorch .pth checkpoint is not an accepted format.
  • The imported model is invoked on demand through Bedrock APIs, with no endpoint to manage, billed for the compute units it uses while active.
  • Limits to remember: embedding models are not supported, and imported models cannot be used with Bedrock batch inference. A model customization (fine-tuning) job is different: it retrains a Bedrock base model on your data rather than importing your weights.

SageMaker AI

  • Prebuilt framework container (PyTorch, TensorFlow, scikit-learn, XGBoost, Hugging Face): package the weights and an inference.py with any custom loading or pre/post-processing (model_fn, input_fn, predict_fn, output_fn) in model.tar.gz, upload it to S3 and deploy. This is the least-effort path for standard frameworks.
  • Custom container (bring your own): an image in Amazon ECR whose web server answers /invocations and /ping on port 8080. Use it for unusual runtimes or dependencies (building containers is Task 3.2).
  • LMI or Triton containers for large language models and multi-framework GPU serving; embedding models that Bedrock cannot import are hosted here.
  • The SageMaker Model Registry catalogs and approves model versions; it does not deploy anything by itself.

How do you deploy and configure AI agents and their tools?

Deploy an agent either as a configured Amazon Bedrock Agents agent, where Bedrock runs the reasoning loop, or as your own agent code hosted on Amazon Bedrock AgentCore Runtime; then give it tools and connect it to other agents with standard protocols.

Hosting choiceYou provideFits
Amazon Bedrock AgentsInstructions, a model, action groups, knowledge bases, guardrails, all set through configurationTeams that do not want to write or host orchestration code
Amazon Bedrock AgentCore RuntimeAgent code in any framework (LangGraph, Strands Agents, CrewAI and others)Existing framework agents; serverless scaling, isolated per-user sessions and long-running sessions (hours)
Self-managed (Lambda, ECS, EKS)Code plus hosting, scaling and isolationOnly when a constraint demands it; Lambda stops at 15 minutes

Action groups: how a Bedrock agent uses tools

  • An action group defines operations with an OpenAPI schema or function details. A Lambda function executes the call, and the agent uses the live result in its reply. Knowledge bases only supply documents, guardrails only filter content, and text in the instructions cannot make network calls.
  • Return of control (RETURN_CONTROL): the agent gathers the parameters and returns them in the InvokeAgent response; the application runs the action with its own context (session, fraud checks) and sends the result back in the session state. No Lambda function is needed.
  • User confirmation on a function makes the agent present the exact proposed call for the user to confirm or deny before it runs; AWS recommends it as a safeguard against prompt injection. A prompt instruction to "ask first" is not enforced.
  • To reach a private API (for example behind an internal load balancer), attach the action group's Lambda function to VPC subnets that can route to it. The agent itself never calls the API, so a VPC endpoint for Bedrock (which serves traffic into Bedrock) does not help.

Agent communication protocols

ProtocolConnectsUse it for
Model Context Protocol (MCP)Agent to toolsA server advertises tools with input schemas that any compliant agent can discover and call. Amazon Bedrock AgentCore Gateway turns existing REST APIs and Lambda functions into MCP tools.
Agent2Agent (A2A)Agent to agentOne agent delegates a task to another agent, possibly built by another team in another framework, and receives the result

Agent memory and state management and the infrastructure behind agentic workflows are covered in Task 3.2.

How do you configure retrieval and reranking in a RAG system?

Retrieval Augmented Generation (RAG) quality depends on retrieving the right chunks and passing only the best of them to the model; diagnose each symptom and change the one retrieval setting that addresses it. Creating knowledge bases and vector stores is Task 3.2; this task is about how queries are run.

SymptomSettingWhy
The right chunk is retrieved but ranked just below the cut-offRetrieve a wider candidate set (several times the number of chunks you will finally pass), apply a reranker model, pass only the top few to the modelA reranker scores each candidate against the query more precisely than vector similarity; passing everything raises token cost and adds noise
Answers draw on documents the user should not see (wrong region, tenant or product line)Metadata filter on the query (for example equals on a region attribute)Filtering scopes retrieval; ranking changes or more results never enforce scope
Comparison or multi-part questions retrieve chunks about only one partQuery decomposition (query transformation, QUERY_DECOMPOSITION) in the orchestration configuration of RetrieveAndGenerate; it is not a Retrieve option, so a pipeline built on Retrieve must split the question itselfSplits the question into sub-queries so each part is retrieved
Not enough relevant context in the promptNumber of resultsMore candidates, at the cost of tokens unless combined with reranking

Where reranking is configured

  • With an Amazon Bedrock knowledge base, set a reranking configuration (a supported reranker model such as Amazon Rerank or Cohere Rerank) on Retrieve or RetrieveAndGenerate.
  • For a custom pipeline that queries a vector store such as Amazon OpenSearch Service directly, call the Amazon Bedrock Rerank API with the candidate documents inline; no extra model has to be hosted. The knowledge base setting does not apply outside a knowledge base.

Retrieve versus RetrieveAndGenerate

Retrieve returns relevant chunks with metadata and source locations, so the application builds its own prompt, adds other data and calls any model. RetrieveAndGenerate (and its streaming form) retrieves and generates a cited answer with a chosen model, which is simplest but gives up control of the prompt. Search type (semantic versus hybrid) and chunking are covered with RAG optimisation in Domain 2.

Tip. Task 3.1 questions describe a model, foundation model, agent or RAG system and its traffic, latency, payload size, team skills and budget, then ask which deployment, configuration or setting fits. One constraint, usually stated mid-scenario, decides between options that would all work: the caller needs a synchronous answer, nothing may be paid while idle, the instance type is fixed, no capacity commitment, no extra hosted model, no Lambda functions, data must stay in one geography, or every service must use existing Kubernetes tooling. Expect near-twin options that differ in one component (serverless versus asynchronous inference, multi-model versus multi-container endpoints, Inferentia versus Trainium, geographic versus global inference profiles, MCP versus A2A, Retrieve versus RetrieveAndGenerate, a knowledge base reranking setting versus the Rerank API), memory arithmetic for GPU sizing, and multi-response items where two settings fix two separate symptoms.

Key takeaways
  • Synchronous plus idle gaps plus acceptable cold start means serverless inference; large payloads or minutes of processing with scale-to-zero means asynchronous inference; a scheduled dataset means batch transform.
  • Asynchronous requests pass the payload as an S3 InputLocation, and success and error SNS topics in AsyncInferenceConfig report completion without polling.
  • Batch transform needs SplitType Line with MultiRecord batching for big files, parallelizes by file, and filters with InputFilter, JoinSource and OutputFilter in that order.
  • Inferentia2 is the low-cost inference accelerator only for Neuron-compatible models, Graviton needs an arm64 image, and GPU instances should be right-sized to the model's memory.
  • Multi-model endpoints share one container across many similar models (Triton on GPUs); multi-container Direct mode hosts different frameworks; serial pipelines chain containers; inference components give each model its own resources and scaling.
  • Large models fit only when weights plus KV cache fit the usable GPU memory: shard with tensor parallelism across every GPU, or quantize when the instance is fixed.
  • Bedrock on-demand bills per token, cross-Region inference profiles absorb peaks without commitment, Provisioned Throughput is a commitment, batch inference serves offline jobs, and Marketplace models run on endpoints you size.
  • Custom Model Import takes supported LLM architectures in Hugging Face safetensors format and serves them on demand, but not embedding models and not with batch inference.
  • Bedrock Agents act through action groups (Lambda or return of control, with optional user confirmation); AgentCore Runtime hosts framework agents; MCP connects agents to tools and A2A connects agents to agents.
  • Reranking fixes good chunks ranked too low, metadata filters fix wrong-scope answers, and query decomposition fixes one-sided multi-part answers.

Frequently asked questions

What is the difference between SageMaker AI serverless inference and asynchronous inference?

Both SageMaker AI serverless inference and asynchronous inference cost nothing while idle, but they serve different callers. Serverless inference answers synchronously in the same request, runs on CPU with a memory size you choose, and suits intermittent traffic where a cold start on the first call is acceptable. Asynchronous inference queues requests whose payload sits in Amazon S3, supports payloads up to 1 GB, processing of up to about an hour and GPU instances, writes results to S3 and can notify through Amazon SNS. Choose serverless for a waiting caller and asynchronous for large or slow requests.

When should I use a SageMaker AI multi-model endpoint instead of a multi-container endpoint?

Use a SageMaker AI multi-model endpoint when you have many models (hundreds or thousands) that share one serving container and are each called infrequently; the endpoint loads each model from Amazon S3 on first use and unloads idle ones. Use a multi-container endpoint in Direct mode when a handful of models need different containers, such as TensorFlow, PyTorch and scikit-learn, and each request targets one container by hostname. If every request must pass through several containers in order, use a serial inference pipeline instead.

What are SageMaker AI inference components?

SageMaker AI inference components let several models share one endpoint while each model declares its own compute requirements, such as the number of accelerators and minimum memory, and scales its own copy count with Application Auto Scaling, including down to zero. Managed instance scaling on the endpoint's production variant adds instances when new copies cannot be placed and removes them afterwards. They are the standard way to pack several foundation models onto multi-GPU instances.

How do I deploy a model trained outside AWS on Amazon Bedrock?

Use Amazon Bedrock Custom Model Import. Upload the model in Hugging Face format to Amazon S3, with weights as safetensors files plus config.json and tokenizer files, and run an import job; the model must use a supported architecture such as Llama or Mistral. The imported model is then invoked on demand through Bedrock APIs without managing an endpoint. Embedding models are not supported and imported models cannot be used with Bedrock batch inference, so host those on SageMaker AI instead.

When should I use Bedrock Provisioned Throughput instead of a cross-Region inference profile?

Use Amazon Bedrock Provisioned Throughput when traffic is steady and heavy enough to justify reserving model units, or when a customized model requires it. Use a cross-Region inference profile when on-demand requests are throttled at peaks and you do not want any capacity commitment: invoking the profile ID routes requests across Regions, and a geographic profile such as US or EU keeps processing within that geography while a global profile does not.

What is the difference between MCP and A2A for AI agents?

The Model Context Protocol (MCP) connects an agent to tools: an MCP server advertises tools with input schemas that any compliant agent can discover and call, and Amazon Bedrock AgentCore Gateway can expose existing APIs and Lambda functions as MCP tools. The Agent2Agent (A2A) protocol connects agents to each other, so one agent can delegate a task to another agent, even one built in a different framework, and receive the result.

How does a Bedrock agent call an external API?

An Amazon Bedrock agent calls external systems through an action group, which defines the operations with an OpenAPI schema or function details. Either a Lambda function executes the call and returns the result to the agent, or the action group returns control so the calling application runs the action itself and sends back the result. User confirmation can be enabled so the user approves each proposed call before it runs, and a Lambda function attached to VPC subnets can reach private APIs.

When should I add a reranker to a RAG system?

Add a reranker when the chunk that answers a question is being retrieved but ranked too low to reach the model. Retrieve a wider candidate set, let a reranker model score each candidate against the query, and pass only the top few to the model, which improves accuracy without paying for a long prompt. Amazon Bedrock knowledge bases support a reranking configuration, and the Bedrock Rerank API serves custom pipelines that query a vector store directly.

Source

This lesson covers the "Deployment and Orchestration of ML and AI Workflows" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.

Test yourself on this topic
Practice questions with full explanations.
Practice now

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?