Choosing ML, Foundation Model and RAG Approaches on AWS (MLA-C02)
Choosing a modeling approach means deciding, before any training starts, whether a business problem is best solved by a pre-trained AWS AI service, a traditional ML algorithm, a foundation model in Amazon Bedrock, a customized FM, or a Retrieval Augmented Generation (RAG) system, and then picking the specific model, customization method, vector store and inference option that meet the requirements at an acceptable cost. Task 2.1 of the MLA-C02 exam tests these decisions through scenarios in which one constraint, such as no labeled data, a published-weights rule, exact identifiers, data residency or idle time, rules out the tempting choice. This lesson works through each decision with the trade-offs that settle it.
On this page10 sections
- How do you choose between traditional ML, an AWS AI service, a foundation model or a custom model?
- Which algorithm or model family fits the problem, and how do interpretability and latency change the choice?
- How do you shortlist and select a foundation model in Amazon Bedrock?
- How do you choose an embedding model?
- When should you use RAG and when should you fine-tune a foundation model?
- Which Amazon Bedrock customization method fits the data you have?
- What is the difference between parameter-efficient fine-tuning (LoRA) and full fine-tuning?
- How do you choose a RAG architecture pattern and vector store?
- How do you trade off performance, training time, latency and cost?
- Which AWS AI service solves which business problem?
- Choose between an AWS AI service, traditional ML, a foundation model and a custom model for a business problem.
- Select an algorithm or model family by problem type, labels, interpretability and latency.
- Shortlist Amazon Bedrock foundation models and embedding models by modality, context window, language and measured performance.
- Decide between RAG and fine-tuning, and pick a customization method (supervised, reinforcement, distillation, continued pre-training, LoRA) from the data available.
- Select a RAG retrieval pattern and a knowledge base vector store from the question types and mandatory features.
- Assess performance, training time, latency and cost trade-offs for ML training, FM inference and custom model hosting.
- Apply Amazon Textract, Rekognition, Comprehend and Transcribe to document, image, text and speech problems.
How do you choose between traditional ML, an AWS AI service, a foundation model or a custom model?
Choose the least custom approach that meets the requirement: a pre-trained AWS AI service when it already does the exact task, a foundation model (FM) when the output is open-ended language or multimodal reasoning, a traditional ML model when you have labeled tabular or high-volume data and need a cheap, fast, repeatable prediction, and a fully custom model only when nothing managed fits. Each step down that list adds control and effort.
| Approach | Fits when | Cost of choosing it |
|---|---|---|
| AWS AI service (Comprehend, Textract, Rekognition, Transcribe) | The task is a standard one the service was pre-trained for (sentiment, invoice fields, unsafe images, speech to text); no labels or ML staff | Least effort; limited control over the model |
| Pre-trained model (SageMaker JumpStart) | A known architecture fits and you want to fine-tune or host it yourself, including export to the edge | You run training and hosting |
| Foundation model (Amazon Bedrock) | Generation, summarization, extraction from free text, chat, reasoning over documents | Pay per token on every request; can be slow and expensive at very high volume |
| Traditional ML (SageMaker AI built-in algorithms or your own code) | Labeled tabular data, fixed classes, forecasts, anomaly scores, millions of predictions a day | You need data and training, but inference is cheap and fast |
| Custom model architecture | Nothing above meets accuracy, interpretability or architecture needs | Most effort and operational ownership |
The common trap is reaching for an FM because it is the modern option. A recurring score for every customer, learned from labeled historical columns, is a supervised tabular problem: a gradient-boosted tree such as XGBoost learns from the labels and scores large populations cheaply, whereas prompting an FM per row is slow, costly and ignores the labels. The same logic applies to high-volume text classification into a fixed set of categories when a large labeled set already exists: a small supervised text classifier (for example the SageMaker AI BlazingText algorithm in supervised mode, or a custom text classification model built without code in Amazon SageMaker Canvas) costs a fraction of per-token FM inference, even discounted batch inference, because the FM bill repeats for every item on every run. Clustering (k-means) is not a substitute for classification: it finds groups but cannot map them to your fixed business categories.
Conversely, when a pre-trained AI service already returns exactly what is asked (for example the four sentiment classes positive, negative, neutral and mixed from Amazon Comprehend), any option that trains or fine-tunes a model is more effort and needs labels you may not have.
Hosting is part of build versus buy
Host a custom model where its architecture, its runtime dependencies and its traffic pattern fit: Amazon Bedrock Custom Model Import for supported open architectures called on demand, a SageMaker AI real-time endpoint when you need a custom container or steady GPU capacity, and SageMaker AI Serverless Inference only for small models that can run on CPU.
| Option | Fits | Rules it out |
|---|---|---|
| Bedrock Custom Model Import | Fine-tuned weights of a supported architecture (for example Llama, in Hugging Face format in Amazon S3); invoked through the Bedrock runtime and billed for use, so idle time costs nothing for inference | Modified or unsupported architectures; it imports weights only, so no extra libraries run alongside the model |
| SageMaker AI real-time endpoint | Custom containers with your own libraries, GPU instances, steady traffic and latency targets | You pay for instances while idle |
| SageMaker AI Serverless Inference | Small models with intermittent traffic; scales to zero | No GPU support and limited memory, so not large models |
| Bedrock Provisioned Throughput on a catalog model | Committed capacity for that base model | Does not include your custom weights |
Deployment mechanics (endpoint types, inference components, auto scaling) are Task 3.1; here the point is that hosting is part of the build-versus-buy decision.
Which algorithm or model family fits the problem, and how do interpretability and latency change the choice?
Match the algorithm to the problem type and to whether you have labels, then let interpretability and latency requirements narrow the field. Accuracy is only one criterion: a model that is marginally more accurate but cannot satisfy a published-explanation rule or a latency budget is the wrong model.
| Problem | Typical choice | Notes |
|---|---|---|
| Tabular classification or regression with labels | XGBoost (gradient-boosted trees) | Strong default for accuracy on tabular data |
| Tabular prediction that must be explained with fixed weights | Linear Learner (linear or logistic regression) | One coefficient per feature, the same for every record |
| Anomaly scores without labeled failures | Random Cut Forest | Unsupervised; scores each point by how unusual it is |
| Forecasts for many related time series, including new items with short history | DeepAR | One recurrent model across all series; outputs probabilistic forecasts (quantiles), useful for safety stock |
| Grouping unlabeled records | k-means | Clusters only; does not classify or score anomalies |
| Classification by similar labeled cases | k-nearest neighbors | Needs labeled examples |
| Images, audio, free text at scale | Deep learning (often a pre-trained network) | Highest capacity; least interpretable |
Interpretability is a requirement, not a preference
Different explanation tools answer different questions, and regulators often require a specific one. A linear model's coefficients are global and fixed: each feature raises or lowers the score by the same amount for everyone. SHAP values (for example from Amazon SageMaker Clarify) are per-prediction attributions, so the contribution of a feature varies from applicant to applicant. A global feature-importance ranking from a tree model says which features matter most but gives neither the direction nor the size of their effect. Monotonic constraints in XGBoost guarantee the direction of a feature's effect but not a constant size. Match the tool to the wording of the rule: a requirement for one fixed, published weight per feature is met only by a linear model, and a small accuracy gap is the accepted price of meeting a mandatory rule.
Solution templates and latency
Amazon SageMaker JumpStart offers pre-trained models and solution templates that package a common use case end to end, which shortens time to a working baseline. For latency, smaller and simpler models (linear models, shallow trees) answer in milliseconds on CPU; large deep networks and FMs need GPUs and take longer. Domain-specific performance should be measured on your own data rather than assumed from general benchmarks.
How do you shortlist and select a foundation model in Amazon Bedrock?
Select an FM in two passes: first apply hard filters that rule models in or out, then compare the survivors on your own task for quality, latency and cost. Comparing price or benchmark scores before checking the hard filters wastes time on models that cannot do the job.
Hard filters (pass or fail)
- Input and output modality: if users send photos with text, the model must accept image input. Check the input side specifically: other capabilities a model has, on the output side or in how it delivers responses, do not substitute for accepting the input type you need.
- Context window: the maximum input tokens per request. If a long document must be read whole in a single request, the context window decides eligibility before anything else. The maximum output tokens is a separate limit on the length of the response.
- Languages the model supports well.
- Features the application needs, such as tool use, streaming, or support for Bedrock customization methods.
- Regional availability where data residency rules apply.
Comparing the survivors
Run a representative set of your real inputs through each candidate and compare output quality, latency and token cost on the same set. General-purpose benchmark scores and model size are proxies that may not reflect your task. Amazon Bedrock provides model evaluation jobs for doing this comparison systematically; how to run and score them is covered in Task 2.3.
How do you choose an embedding model?
Choose an embedding model by the modality of what you need to search, the languages involved, and the vector size you can afford; the same model must embed both the documents and the queries, because retrieval compares vectors in one shared space. Only embedding models produce vectors; a text generation model returns text, not an embedding.
| Model | Use it for | Watch out for |
|---|---|---|
| Amazon Titan Text Embeddings V2 | Text retrieval; output size selectable as 256, 512 or 1,024 dimensions (1,024 is the default) | Accepts many languages but is optimized for English; AWS states cross-language queries (asking in one language about text in another) give sub-optimal results |
| Amazon Titan Multimodal Embeddings G1 | Images and text in one vector space: photo-to-photo similarity and text-to-image search | Only helps if the images themselves are embedded; text input is English only |
| Cohere Embed English v3 | English text retrieval; also accepts image input | English only |
| Cohere Embed Multilingual v3 | Multilingual and cross-language text retrieval; also accepts image input | Up to 512 tokens per text input; one image per request |
The table covers common choices, not the whole catalog: newer models such as Cohere Embed v4 are also available in Amazon Bedrock, so check the current model list and each model's modalities, languages and input limits.
Two decisions recur. First, for visual search, embed the product images with a multimodal model; embedding only titles and descriptions with the same model means every query is still matched against text. Reducing images to text first (captions or detected labels) and embedding that text discards the visual features that make two images look alike. Second, vector storage and index memory scale with the number of chunks times the number of dimensions. When storage dominates cost and a small accuracy loss is acceptable, re-embedding at a lower output dimension (256 instead of 1,024 is a 75% reduction per vector) keeps the same model and tokenization. Changing the dimension or the model always means re-embedding the whole corpus; tuning chunking and embeddings in detail belongs to Task 2.2.
When should you use RAG and when should you fine-tune a foundation model?
Use Retrieval Augmented Generation (RAG) when the problem is knowledge, and fine-tuning when the problem is behaviour. RAG retrieves relevant passages at query time and gives them to the model, so it handles facts that change, answers that must cite a source, and per-user restrictions on which documents can be used. Fine-tuning changes the model's weights from examples, so it handles tone, fixed output structure, and narrow task conventions such as a proprietary labeling taxonomy.
| Requirement | RAG (for example an Amazon Bedrock knowledge base) | Fine-tuning |
|---|---|---|
| Facts change weekly or monthly | Yes: re-sync the data source, no training | No: each run bakes in a snapshot that goes stale |
| Cite the source document or clause | Yes: knowledge bases return citations | No: the weights cannot point to a source |
| Restrict answers to documents a user may see | Yes: metadata filtering at query time | No |
| Consistent tone, style or section layout | Weak: layout rules stored as documents are just more prompt text | Yes, given enough good examples |
| Classify into a company-specific label set | Weak | Yes |
The two combine. When an application needs current, citable facts and a behaviour that prompting alone fails to enforce, use retrieval for the facts and supervised fine-tuning on approved examples for the behaviour. Context copied into a static prompt by hand is not retrieval: it must be maintained manually, goes stale and carries no citations.
Which Amazon Bedrock customization method fits the data you have?
Pick the customization method by the data and signal you actually have, because each method needs a different input. The question to ask first is not "which method is best" but "do we have labeled pairs, a grader, a good teacher model, or only raw text?"
| Method | Input it needs | What it changes |
|---|---|---|
| Supervised fine-tuning | Labeled prompt-and-response pairs | Task behaviour: format, tone, labeling conventions |
| Reinforcement fine-tuning | Training prompts plus a reward function (for example an AWS Lambda function that grades each output) | Improves outputs that can be scored automatically even when no reference answers exist, such as outputs checked by an automated validator |
| Model distillation | Use-case prompts (or invocation logs) and a teacher model whose accuracy is acceptable | A smaller, faster, cheaper student that reproduces the teacher on your use case; the teacher generates the responses, so no human labels are needed |
| Continued pre-training | Large volumes of unlabeled domain text | Domain vocabulary and language, before any task-specific work (SageMaker JumpStart calls the equivalent for its open models domain adaptation fine-tuning) |
Limits matter as much as fit. Distillation cannot exceed its teacher: if no available model performs the task well, there is nothing good to distil, whereas any automatic way of grading outputs can drive reinforcement fine-tuning. Supervised fine-tuning needs examples of the desired outputs; reference documents on their own are material for continued pre-training or retrieval, not labeled pairs. Capacity options change throughput, not what the model has learned. Prompt engineering should always be tried first; the mechanics of prompting and of running training jobs are Task 2.2.
What is the difference between parameter-efficient fine-tuning (LoRA) and full fine-tuning?
Parameter-efficient fine-tuning (PEFT) such as LoRA freezes the base model and trains small low-rank adapter matrices, while full-parameter fine-tuning updates every weight. LoRA therefore needs far less GPU memory and compute, trains faster, and produces an adapter that is a small fraction of the model's size; full fine-tuning can extract slightly more task performance but needs much more compute and produces a complete copy of the model per variant.
| Concern | LoRA adapters | Full fine-tuning |
|---|---|---|
| Training compute and memory | Low | High |
| Storage per variant | Small adapter | Full model copy |
| Many variants refreshed independently | Retrain one adapter on its own | Retrain and store a full model each time |
| Hosting many variants | Many adapters can share one copy of the base model | Each variant needs its own hosting |
The hosting saving is what makes LoRA decisive for many low-traffic variants. Amazon SageMaker AI can serve multiple LoRA adapters as inference components on a single endpoint that holds one copy of the base model, so many variants share GPU capacity. The saving depends on keeping adapters separate at serving time: any design that turns each variant back into a full-size model gives it up, and folding all variants into a single model gives up independent updates.
How do you choose a RAG architecture pattern and vector store?
Choose the RAG pattern by the shape of the questions and of the data: plain semantic vector search answers descriptive questions, but exact identifiers, multi-hop relationships, computed aggregates and access restrictions each need a different pattern. Diagnosing which kind of question fails tells you which pattern to add.
| Symptom or requirement | Pattern | Why |
|---|---|---|
| Exact part numbers, SKUs or fault codes retrieve similar-sounding wrong items | Hybrid search (overrideSearchType set to HYBRID) | Adds keyword matching to vector search; embeddings are weak at exact strings |
| Answers need facts linked across several documents (supplier to component to recall) | GraphRAG with Amazon Neptune Analytics | Extracts entities and relationships and uses them during retrieval |
| Exact totals and comparisons over warehouse tables | Bedrock knowledge base connected to a structured data store such as Amazon Redshift | Converts the question into SQL, so aggregates are computed, not guessed from chunks |
| Users may only see certain documents | Metadata filtering at query time | Restricts retrieval to permitted documents |
| A question bundles several sub-questions | Query decomposition | Splits it into sub-queries; does not change how relevance is computed |
Hybrid search in Amazon Bedrock Knowledge Bases is supported only for Amazon RDS (Aurora PostgreSQL), Amazon OpenSearch Serverless and MongoDB vector stores that include a filterable text field; other stores fall back to semantic search. Raising the number of results returned does not fix ranking that is computed the wrong way, and a reranker can only reorder chunks that retrieval actually returned. Text retrieval over chunked copies of table data, or training on such copies, cannot compute reliable aggregates; only an engine that runs the query over the rows can. Amazon Kendra is a document search service: it retrieves documents (and can honour document access-control lists) rather than computing results. Tuning chunk size, reranking and retrieval settings in detail is covered in Tasks 2.2 and 3.1.
Choosing the vector store behind a knowledge base
Choose the vector store by the features you must have first (hybrid search, graph retrieval, SQL joins) and by cost and scale second. A cheaper store that lacks a mandatory feature is not an option, however attractive its price.
| Store | Best fit | Limits |
|---|---|---|
| Amazon OpenSearch Serverless | The quick-create default; supports hybrid search; steady query traffic | Capacity-based billing, so a lightly used index still costs money |
| Amazon S3 Vectors (vector bucket and vector index) | Very large, infrequently queried collections where sub-second latency is acceptable; pay for storage and queries with nothing to provision | Semantic search only, no hybrid search |
| Amazon Aurora PostgreSQL with pgvector | Vectors stored beside relational data you already run, so SQL can join them; supports hybrid search | You operate the database |
| Amazon Neptune Analytics | GraphRAG over entities and relationships | A graph engine, chosen for multi-hop questions rather than cost |
Read the scenario for the deciding constraint. When hybrid search is mandatory, pick a store that supports it even if query volume is low. When a collection is very large, rarely queried and purely semantic, S3 Vectors usually costs least. When vectors must be joined with records already held in Aurora PostgreSQL, add pgvector to that database rather than copying data to a new store. Creating and indexing the knowledge base itself is Task 3.2.
How do you trade off performance, training time, latency and cost?
Trade-offs differ for training and for inference: for ML training, reduce time and cost by starting from pre-trained weights, paying less for the same compute, and choosing the simplest model whose accuracy meets the requirement. Faster hardware or more instances finish sooner but rarely cost less, and a small accuracy gain seldom justifies a much more expensive model.
- Transfer learning: fine-tuning a pre-trained network (for example an image classifier from SageMaker JumpStart) on a modest labeled set is faster than training from random weights and overfits less, because the network already knows general features. Distributed training or Spot pricing changes how fast or at what price a from-scratch job runs, not how much work it needs. If the model must be exported (for example to edge devices), a managed service whose models are only served by that service, such as Amazon Rekognition Custom Labels, does not qualify.
- Managed Spot training: runs SageMaker AI training on spare capacity at a large discount. Set a maximum wait time longer than the maximum run time so interruptions are tolerated, and save checkpoints to Amazon S3 so the training script resumes from the latest checkpoint instead of restarting. Spot without checkpoints risks losing hours of progress. Use it when the deadline has slack.
- Warm pools keep training instances provisioned between jobs to cut start-up latency; you pay for the retained instances, so they reduce waiting, not cost.
- Right-sizing: a larger instance may finish sooner but is billed at a higher rate; measure before assuming it is cheaper.
For foundation model inference, the levers are different and are covered below. Techniques inside the training job, such as early stopping, distributed training and hyperparameter tuning, belong to Task 2.2.
Foundation model performance, latency and cost
Lower FM cost and latency by matching each workload to the right lever: a smaller or distilled model where quality allows, prompt caching for repeated prompt prefixes, batch inference for non-interactive bulk jobs, Provisioned Throughput for steady high load, and cross-Region inference or routing for throughput and mixed difficulty. Each lever solves one problem and creates a constraint, so read the scenario's constraint before choosing.
| Lever | Solves | Does not fit when |
|---|---|---|
| Smaller model in the same family | Cost and latency on easy tasks | Quality on hard requests must not drop |
| Intelligent prompt routing | Mixed easy and hard traffic: routes each prompt between two models in the same family by predicted quality | Prompts are not mainly English (it is optimized for English) |
| Model distillation | A cheaper student that matches the large model on specific intents | No teacher is good enough |
| Prompt caching | A long identical prefix (instructions, rules) repeated in every request; lowers cost and latency | Prompts share no repeated prefix |
| Batch inference | Large, non-interactive jobs with hours of slack, at a discounted price | Users wait for the answer |
| Provisioned Throughput | Steady, high traffic that is throttled on demand; reserved model units in one Region | Traffic is sporadic, or you hoped it would lower per-request cost or latency |
| Cross-Region inference profile | Throttling at peaks, by routing requests to other Regions | Inference must stay in a single Region |
Retrieved context counts as input tokens exactly like typed prompt text, so relocating repeated text does not shrink a request; caching lowers the cost of a repeated prefix instead. Moving all traffic to a cheaper model trades away quality on the hard requests, whereas routing and distillation aim only at the share a smaller model can handle.
Which AWS AI service solves which business problem?
Use Amazon Textract for documents, Amazon Rekognition for images and video, Amazon Transcribe for speech, and Amazon Comprehend for text; each offers pre-trained operations for common needs and customization only where the built-in model cannot know your vocabulary or labels.
Amazon Textract
| Operation | Returns |
|---|---|
| DetectDocumentText | Raw lines and words, no structure |
| AnalyzeDocument with FORMS / TABLES / QUERIES | Key-value pairs, table cells, or answers to natural-language questions about the page; keys are whatever each layout prints |
| AnalyzeExpense | Invoices and receipts: normalized fields (vendor, total, due date) plus line items across any vendor layout |
| AnalyzeID | Normalized fields from identity documents such as driver's licenses and passports |
Use AnalyzeExpense when invoices or receipts arrive in many layouts, because generic FORMS output keeps each layout's own labels and TABLES gives line items without normalized header fields. Mixed document sets combine operations, each on the documents it is built for: AnalyzeDocument for general forms and tables, AnalyzeID for identity documents. Rekognition text detection reads text in photos and video frames but returns no document structure or identity fields.
Amazon Rekognition
DetectModerationLabels returns a hierarchy of unsafe-content labels (explicit, violent and so on) with confidence scores, which can drive automatic blocking or send borderline images to human review. DetectLabels identifies general objects and scenes and is not a moderation model. Custom Labels trains a model on your own labeled images for business-specific objects; it needs labeled data, its models are served by Rekognition, and AWS does not position it for unsafe-content detection. To tune moderation to your own policy, Rekognition Custom Moderation trains an adapter on your annotated images that improves DetectModerationLabels accuracy.
Amazon Transcribe and Amazon Comprehend
Transcribe converts audio to text. PII redaction removes sensitive data such as card and Social Security numbers from transcripts, and a custom vocabulary improves recognition of product names and jargon. Comprehend works only on text, so audio must be transcribed first. Its pre-trained APIs detect sentiment (positive, negative, neutral, mixed), built-in entities, PII and key phrases; asynchronous jobs process large document collections, and per-document size limits differ by API. When the built-in entity types do not cover your terms, train a custom entity recognizer (from an entity list or annotations) to extract mentions; a custom classifier is the near-twin that assigns a label to a whole document instead of extracting mentions.
Some older purpose-built services are closed to new customers (for example Amazon Forecast); for a new forecasting workload use SageMaker AI, such as the DeepAR algorithm.
Tip. Task 2.1 questions are scenarios that describe a business problem and its data, then ask which approach, model, method or configuration fits. The deciding detail is usually one constraint in the middle of the story: no labeled data or no reward function, a regulator's explanation rule, exact identifiers in queries, questions that link several documents, a single-Region residency rule, long idle periods, a modified model architecture, export to edge devices, or a fixed cost ceiling at high volume. Distractors are workable-sounding near-twins that miss that one constraint (the same embedding model applied to the wrong input, a cheaper vector store without hybrid search, Spot training without checkpoints, a custom classifier instead of an entity recognizer), and the modern-sounding option, such as a foundation model for tabular prediction, is often wrong. Multi-response items commonly pair two tools that each fix a separate part of the problem.
- Use the least custom approach that works: AI service, then FM or traditional ML, then a custom model; FMs are a poor fit for high-volume tabular or fixed-class prediction when labels exist.
- Interpretability requirements decide the model: fixed published weights mean a linear model; SHAP values vary per prediction, importance rankings lack direction, monotonic constraints fix direction only.
- Filter FMs on hard requirements (modality, context window, languages, Region) first, then compare candidates on your own prompts for quality, latency and cost.
- The same embedding model must embed documents and queries; use a multimodal model on the images for visual search and a multilingual model for cross-language retrieval; fewer dimensions means less storage.
- RAG handles changing facts, citations and access control; fine-tuning handles tone, layout and labeling behaviour; many systems need both.
- Customization follows the data: labeled pairs for supervised fine-tuning, a reward function for reinforcement fine-tuning, a good teacher and prompts for distillation, unlabeled text for continued pre-training.
- LoRA adapters cut training cost and let many variants share one base model on a SageMaker AI endpoint.
- Hybrid search fixes exact identifiers, GraphRAG fixes multi-hop questions, structured data stores compute exact aggregates; a missing mandatory feature rules out a cheaper vector store.
- Prompt caching, batch inference, Provisioned Throughput, prompt routing and cross-Region inference each solve one cost or capacity problem and carry one constraint.
- Textract AnalyzeExpense for invoices, AnalyzeID for identity documents, Rekognition DetectModerationLabels for unsafe images, Transcribe for speech with PII redaction and custom vocabulary, Comprehend for text.
Frequently asked questions
When should I use RAG instead of fine-tuning a foundation model?
Use RAG when answers depend on facts that change, must cite their source, or must be limited to documents a user may see, because retrieval happens at query time and a knowledge base can be re-synced without training. Use fine-tuning when you need to change the model's behaviour, such as tone, a fixed output layout or a company-specific label set, and you have good examples. When you need both current facts and a consistent behaviour, combine them.
What is the difference between model distillation and fine-tuning in Amazon Bedrock?
Supervised fine-tuning trains a model on prompt-and-response pairs you supply. Model distillation needs only your use-case prompts (or invocation logs) and a teacher model with acceptable accuracy: the teacher generates the responses and a smaller student model is fine-tuned on them, so you get lower cost and latency without writing labels. A distilled student cannot be better than its teacher.
When is reinforcement fine-tuning the right choice?
Reinforcement fine-tuning fits when you cannot write reference answers but can score an output automatically, for example by running automated checks against it. You supply training prompts and a reward function, such as an AWS Lambda function, and the model learns from the scores.
Which vector store should I use for an Amazon Bedrock knowledge base?
Start from mandatory features. Hybrid keyword-plus-semantic search needs a store that supports it, such as Amazon OpenSearch Serverless or Aurora PostgreSQL with pgvector. GraphRAG needs Amazon Neptune Analytics. Joining vectors with relational data points to pgvector in your existing Aurora PostgreSQL cluster. For very large, rarely queried, purely semantic collections, Amazon S3 Vectors is usually the lowest-cost option.
How do I reduce the cost of a foundation model workload in Amazon Bedrock without changing models?
Enable prompt caching when requests repeat a long identical prefix, and move large non-interactive jobs to batch inference, which is priced at a discount. Provisioned Throughput suits steady high traffic that needs guaranteed capacity rather than a lower price, and cross-Region inference helps with throttling but sends requests to other Regions.
Is LoRA better than full fine-tuning?
LoRA is usually the better trade-off when compute is limited or you need several variants: it trains small adapters on a frozen base model, so training is cheaper, each variant is small, and SageMaker AI can host many adapters on one endpoint over a shared base model. Full fine-tuning updates every weight and costs more to train, store and host.
Which Amazon Textract API should I use for invoices?
Use AnalyzeExpense. It returns normalized fields such as vendor name, invoice total and due date, plus line items, across different vendor layouts, so you do not maintain per-vendor templates. AnalyzeDocument with FORMS returns each layout's own labels, and AnalyzeID is for identity documents.
How do I choose a foundation model in Amazon Bedrock?
First rule models in or out on hard requirements: input modality (for example image input), context window, supported languages, needed features and Regional availability. Then run a representative set of your own inputs through the remaining candidates and compare output quality, latency and token cost. General benchmark scores and model size are only proxies.
Source
This lesson covers the "ML Model and Foundation Model (FM) Development" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.
- AWS Certified Machine Learning Engineer – Associate (MLA-C02) exam guide — Amazon Web Services
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.