Data Transformation, Feature Engineering and Pre-processing on AWS (MLA-C02)
Data transformation, feature engineering and pre-processing turn raw data into the inputs a model can learn from: cleaned and reshaped tables, scaled and encoded features, embeddings, chunked documents and correctly formatted training files. On AWS that work runs on AWS Glue, Glue DataBrew, Amazon EMR, SageMaker Data Wrangler, SageMaker Feature Store, AWS Lambda and Amazon Bedrock. Task 1.2 of the MLA-C02 exam covers it for both traditional ML and foundation models: choosing the transformation tool, managing features without leakage, transforming streams, the classic feature-engineering techniques, embeddings and tokenization, preparing documents for RAG, protecting sensitive data, and formatting data for fine-tuning, continued pre-training and distillation.
On this page10 sections
- Which AWS tool should transform the data?
- How do you create and manage features in SageMaker Feature Store?
- How do you transform streaming data with Lambda or Spark?
- When do features need scaling, and how should you fit the scaler?
- How do log transformation, binning, feature splitting and encoding reshape features?
- How do you configure embedding models for text and images?
- What is tokenization, and how do you augment text without changing labels?
- How do you prepare documents for RAG in an Amazon Bedrock knowledge base?
- How do you mask, redact and anonymize sensitive data?
- How do you prepare data for fine-tuning, continued pre-training and distillation?
- Choose between AWS Glue, Glue DataBrew, Spark on Amazon EMR and SageMaker Data Wrangler for a transformation, and productionise a Data Wrangler flow by exporting it.
- Configure SageMaker Feature Store feature groups (record identifier, event time, online and offline stores, InMemory tier, TTL) and build point-in-time correct training sets.
- Decide when a streaming transformation belongs in Lambda and when it needs Spark Structured Streaming or Flink, including tumbling windows and watermarks.
- Apply scaling, standardization, normalization, log transformation, binning, feature splitting and categorical encoding correctly, fitting transforms on training data only.
- Configure embedding models and tokenization consistently, and augment specialised text without changing its labels.
- Prepare documents for Amazon Bedrock Knowledge Bases with the right chunking strategy, metadata files and parser.
- Mask, redact and pseudonymize sensitive data, and format data for Bedrock fine-tuning, continued pre-training and distillation.
Which AWS tool should transform the data?
Choose the transformation tool by who builds the logic, how much infrastructure the team will manage, and whether the work is general data engineering or ML-specific preparation. AWS Glue runs serverless Spark ETL, AWS Glue DataBrew gives non-coders a visual recipe editor, Amazon EMR gives full control of a Spark cluster, and Amazon SageMaker Data Wrangler (now part of SageMaker Canvas) is a visual tool built for preparing ML features. All four can read from and write to Amazon S3, so the deciding factor is almost always a constraint in the scenario rather than the data location.
| Tool | Best fit | What you manage | Watch out for |
|---|---|---|---|
| AWS Glue ETL (Spark) | Code-based, scheduled or event-driven ETL with no servers; incremental runs with job bookmarks | Script, job settings, workers (DPUs) | Less control over Spark internals and native libraries than a cluster you own |
| AWS Glue DataBrew | Analysts who do not write code: visual cleaning and normalising steps saved as a recipe and rerun as a scheduled recipe job | Recipes and jobs only | Tabular data prep, not custom algorithms or streaming |
| Spark on Amazon EMR | Very large or specialised Spark jobs that need custom native libraries, tuned Spark settings, or cost control with Spot Instances | Cluster configuration, scaling, patching (less with EMR Serverless) | Highest operational overhead of the four |
| SageMaker Data Wrangler | Data scientists exploring and transforming data for a model: joins, encoding, scaling, built-in analyses, then export | The flow; exports run as processing jobs | Interactive by design; production runs come from exporting the flow |
Glue job bookmarks
A job bookmark records what a Glue job has already processed, so the next run reads only new files in an S3 prefix instead of reprocessing the full history. Moving a PySpark script from a long-lived EMR cluster to a Glue Spark job with bookmarks removes cluster management and makes the job incremental in one step. Glue bills per DPU-hour for the workers a job uses, and for work that is not time-sensitive, the Flex execution class runs on spare capacity at a lower price.
Lowering EMR cost
When a Spark job tolerates task retries and the team accepts cluster management, EMR can run task nodes on Spot Instances while core nodes stay On-Demand. Core nodes hold HDFS data and the cluster's stability, so losing them hurts; task nodes only compute, so losing one just reschedules its tasks.
Putting a Data Wrangler flow into production
You put a SageMaker Data Wrangler flow into production by exporting it, not by rewriting its logic by hand. A flow (the .flow file) records every join, encoding and scaling step, and Data Wrangler can export it so the same steps run again on new data.
- SageMaker Pipelines: the flow becomes a processing step in a pipeline, which is the right target when the transformation is one stage of a scheduled or automated training workflow.
- SageMaker Feature Store: the output is ingested into a feature group, so other teams reuse the same features for training sets (offline store) and real-time lookups (online store) instead of rerunning the flow on their own copies.
- Amazon S3 through a processing job: a one-off or ad hoc run that writes the transformed dataset to S3.
- Python code: the steps as a script, when you need to embed them in your own code.
Re-implementing a working flow in a new script is the slow, error-prone option: the two versions drift, and the model in production sees features computed differently from the ones it was trained on.
How do you create and manage features in SageMaker Feature Store?
SageMaker Feature Store organises features into feature groups, and every feature group must name a record identifier feature (who the row describes, such as customer_id) and an event time feature (when that version of the values was true, such as updated_at). Each write with a new event time adds a new version of the record. Those two names are what make both fast lookups and time-travel history possible. (How data is ingested into a feature group belongs to Task 1.1; this lesson covers using the features.)
| Online store | Offline store | |
|---|---|---|
| Holds | The latest version of each record | Every version of every record, appended in Amazon S3 |
| Access | GetRecord / BatchGetRecord by record identifier, in milliseconds | Queries through the AWS Glue Data Catalog (for example Amazon Athena) or the SDK dataset builder |
| Used for | Real-time inference | Building training sets, history, audits |
| Options | Standard or InMemory storage type; TtlDuration | Table format (Glue or Apache Iceberg) |
Online store options
The InMemory storage type gives the lowest read latency for workloads where every millisecond of online lookup counts; the Standard tier suits most workloads. TtlDuration on the online store makes a record expire a set time after its event time, so a record that is no longer refreshed stops being served for real-time inference once it passes its expiry. TTL applies to the online store only: the offline store keeps the full history, and expired records are removed from the online store asynchronously rather than at the exact instant.
Point-in-time correct training sets
You build a leak-free training set with a point-in-time correct join: each label row is matched to the feature record version with the latest event time at or before the label's own timestamp. Joining labels to the online store, or to the newest offline values, gives every historical label features from its future, which inflates validation scores and then disappoints in production.
- Keep the offline store enabled from the start, so every version is retained.
- Put labels in a DataFrame with the record identifier and a timestamp for when the label applied (for example label_time or approval time).
- Use the SageMaker Python SDK dataset builder (create_dataset with a point-in-time accurate join) or an equivalent Athena query to join on the event time feature.
The offline store also has system columns such as write_time (when Feature Store wrote the row) and api_invocation_time. These record ingestion, not when the values were true: data loaded in bulk later carries recent ingestion timestamps for old facts, so point-in-time logic must always key on the event time feature you defined.
How do you transform streaming data with Lambda or Spark?
Use AWS Lambda for stateless, per-record transformations of a stream, and Spark Structured Streaming (AWS Glue streaming jobs or Amazon EMR) or Apache Flink (Amazon Managed Service for Apache Flink) for stateful, windowed aggregations that must handle late data. The deciding question is whether a record can be processed on its own.
| Need | Fit |
|---|---|
| Clean, drop or enrich records on a Firehose stream before delivery to S3 | Amazon Data Firehose data transformation with a Lambda function; the function returns every record with its recordId, a result of Ok, Dropped or ProcessingFailed, and the base64-encoded transformed data |
| Enrich each Kinesis or Kafka record and write it to a store such as DynamoDB | Lambda through an event source mapping |
| Short, fixed, non-overlapping per-key aggregates with an existing Lambda consumer | Event source mapping tumbling window (up to 15 minutes); state is carried between invocations per shard, so the partition key should be the entity being aggregated. Batching settings only control how many records an invocation receives; they keep no state between invocations |
| Sliding windows, event-time windows, late events, joins across streams | Spark Structured Streaming or Apache Flink |
Record-level logic such as filtering or enrichment needs the Lambda transformation; Firehose's built-in format conversion (JSON to Parquet or ORC) only changes how records are stored. Tumbling windows only cover fixed, back-to-back intervals and process by arrival, not event time, so a sliding window recomputed every minute or an event that arrives late belongs in Spark or Flink.
Watermarks
In Spark Structured Streaming, a watermark on the event-time column (withWatermark, defined before the windowed aggregation) tells the engine how late an event may arrive and still be counted. Windows older than the watermark are finalised and their state is dropped, which also stops state from growing without limit. Without a watermark, Spark keeps every window open. A trigger interval is a separate setting: it controls how often micro-batches run, not how late data may be.
When do features need scaling, and how should you fit the scaler?
Features need scaling when the algorithm uses distances, dot products or gradient-based optimisation, because a feature measured in large units otherwise dominates one measured in small units. k-nearest neighbors, k-means, support vector machines, linear and logistic regression and neural networks are sensitive to scale; tree-based models such as XGBoost split on thresholds and are largely unaffected.
| Technique | Formula | Use when |
|---|---|---|
| Standardization (z-score) | (x − mean) / standard deviation | Default for distance and gradient methods; result has mean 0, variance 1 |
| Min-max normalization | (x − min) / (max − min) | A bounded range such as 0 to 1 is needed and there are no extreme outliers |
| Robust scaling | (x − median) / IQR | Legitimate outliers must stay in the data; they no longer squeeze everything else into a tiny range |
Fit on training data only
Any fitted transformation (scaler, encoder, imputer) learns its parameters from the training split only, and those saved parameters are then applied unchanged to validation, test and every inference request. Statistics computed anywhere else, whether from the full dataset before splitting or from the data arriving at inference, either leak test information into training or feed the model inputs on a different scale from the one it learned. Removing or capping extreme values is a data-cleaning decision about errors, not a choice of scaler.
How do log transformation, binning, feature splitting and encoding reshape features?
Use a log transformation to compress a long right tail, and binning to let a model that only learns straight-line effects capture a relationship that is non-linear or changes direction. Both reshape a feature; neither removes data.
Log transformation
Taking the log of a right-skewed feature (income, deposits, prices) stops a few huge values from dominating a linear model. log(x) is undefined at 0, so a feature with zeros uses log(1 + x) (log1p). The same idea applies to a skewed target: training on log(price) balances errors across cheap and expensive items, but predictions then come out in log units. Apply the exponential function to convert each prediction back, and compute business metrics such as RMSE on the back-transformed predictions against the actual values; an error measured in log units is not an error in dollars, and exponentiating a log-scale RMSE gives a multiplicative ratio, not a dollar amount.
Binning
Binning (bucketing) groups a numeric feature into ranges. A linear model given outdoor temperature as one number can only learn one slope, so it cannot learn energy demand that is high on cold days, low on mild days and high again on hot days. Binning temperature into ranges and one-hot encoding the bins gives each range its own weight; anything that leaves the feature as a single numeric column still gives it one slope. Quantile binning gives each bin a similar number of records; equal-width binning uses fixed intervals.
Feature splitting and categorical encoding
Feature splitting breaks a compound value into the parts that carry signal, and categorical encoding turns labels into numbers in a way that matches what the labels mean. A raw timestamp string is useless to most models, but its components are not.
- Timestamps: extract hour of day, day of week, month or a weekend flag. Choose components that repeat with the pattern you expect: hour of day and day of week capture daily and weekly cycles, while components that never repeat, or repeat on a cycle unrelated to the behaviour, add noise rather than signal. Convert to the local time zone of the event first when behaviour follows local clocks (meal times, commuting); a UTC hour shifts the pattern for every city.
- Compound strings: split values such as "city, country" or product codes with meaningful segments into separate columns.
| Category type | Example | Encoding |
|---|---|---|
| Nominal (no order) | Payment method, colour | One-hot encoding: one binary column per value |
| Ordinal (natural order) | Shirt size S, M, L, XL | Ordinal encoding: integers in the real order |
Mapping nominal values to integers tells a linear model that "cash" is five times "card" and that the values are evenly spaced, which is false. Any scheme that puts every category on one numeric scale implies an order the data does not have. Very high-cardinality nominal features may need other approaches (such as target or hashing encoding, with enough hash buckets that unrelated values rarely collide; too few buckets merges unrelated categories), but for a handful of unordered values, one-hot encoding is the standard choice.
How do you configure embedding models for text and images?
An embedding model converts text or images into fixed-length numeric vectors whose distances reflect meaning, and configuring one means choosing the model, its output dimension and normalisation, then using exactly the same configuration for documents and queries. On Amazon Bedrock, Amazon Titan Text Embeddings V2 outputs 1,024 dimensions by default and can be set to 512 or 256 to cut storage and search cost for a small accuracy trade-off; it can also return normalised vectors for cosine or dot-product similarity. Multimodal models (such as Amazon Titan Multimodal Embeddings or Amazon Nova multimodal embeddings) map images and text into one shared space, so a typed description or an uploaded photo can be searched against the same image index.
Consistency rules
- The vector index field's dimension must equal the model's output dimension. Changing the dimension means creating a new index with the new dimension, re-embedding every existing document and switching traffic; vectors from different configurations are not comparable.
- Queries must be embedded with the same model (and version and settings) as the documents. Two models with the same dimension still produce unrelated vector spaces, so nearest-neighbour distances between them carry no meaning.
- Text-only models cannot embed images; when the images themselves must be searchable, use a multimodal model.
Choosing which embedding model gives the best retrieval quality, and how retrieval itself is configured, is covered with RAG architecture in Domain 2 and Domain 3.
What is tokenization, and how do you augment text without changing labels?
Tokenization splits text into the units a model actually reads, usually subword pieces, and every limit on model input or output is measured in tokens of that model's own tokenizer. A word can be one token or many: serial numbers, URLs, gene identifiers, stack traces and non-English text break into far more tokens than everyday words, so a word or character count is not a safe proxy. Size inputs by counting tokens with the target model's tokenizer (or the provider's token-count tooling) and leave headroom below the limit.
Domain-specific augmentation
Text augmentation creates new training examples from existing ones, and in a specialised domain it must preserve each example's label. Generic synonym swaps can turn one medical or legal term into another with a different meaning, so the augmented example no longer matches its label. Safer methods:
- Replace terms with equivalents from a curated domain vocabulary or ontology, and have experts review samples.
- Back-translation: translate each example to another language and back (for example with Amazon Translate) to get a paraphrase that keeps the original label, with no model to train or host.
- Template or LLM-generated paraphrases, reviewed before use.
The test for any method is whether the new example is realistic text the model will meet at inference and still carries its original label. Methods designed for numeric feature vectors, or edits applied at random, often fail one or both.
Augmentation as a fix for class imbalance and bias, and validating the resulting data, belongs to Task 1.3.
How do you prepare documents for RAG in an Amazon Bedrock knowledge base?
You prepare documents for RAG by choosing how they are parsed, chunked and tagged with metadata before they are embedded. Chunking splits documents into the passages that are embedded and retrieved, and the right strategy depends on how self-contained your documents are and whether you need precise matching, wide context or both. Amazon Bedrock Knowledge Bases set the chunking strategy when the data source is created; it cannot be changed afterwards, so a new strategy means a new data source.
| Strategy | How it splits | Use when |
|---|---|---|
| Default | Chunks of roughly 300 tokens, respecting sentence boundaries | A reasonable starting point for general documents |
| Fixed-size | Your maximum tokens per chunk plus an overlap percentage | You want predictable chunk sizes; overlap keeps sentences that span a boundary |
| Hierarchical | Small child chunks inside larger parent chunks; search matches children, the parent is returned | You need precise matching and the surrounding context at the same time |
| Semantic | Splits where meaning changes, using maximum tokens, a buffer size (how many neighbouring sentences are grouped when comparing embeddings) and a breakpoint percentile threshold | Topics shift unevenly and fixed boundaries cut ideas in half |
| No chunking | Each file is one chunk | Files are already short, self-contained units, or you pre-split them yourself |
Large fixed chunks add context but blur each embedding across several topics, lowering precision; small chunks match precisely but lose neighbouring conditions. Hierarchical chunking resolves that trade-off. In semantic chunking, a higher breakpoint percentile threshold means only bigger shifts in meaning start a new chunk, giving fewer, larger chunks; lowering it, or lowering maximum tokens, gives smaller chunks with sharper boundaries. Buffer size smooths how similarity is measured around each sentence; it is not a chunk-size control.
Custom chunking
When none of the built-in strategies can reproduce your splitting logic (for example splitting at markers specific to your own document format), set the data source to no chunking and attach a custom transformation Lambda function, which also requires an intermediate S3 location for the files it reads and writes. The knowledge base keeps managing ingestion, embedding and sync.
Metadata and advanced parsing
Metadata lets retrieval filter or prioritise chunks by attributes such as product version, department or date, and advanced parsing makes sure tables, charts and images are extracted correctly before chunking. Both are set up at ingestion time.
Metadata files
For an S3 data source, each document can have a sidecar file named <document name>.metadata.json in the same location (for travel.pdf, travel.pdf.metadata.json), containing a metadataAttributes object such as {"department": "finance"}. Ingestion reads it on the next sync of the data source; until then, filters cannot see it. At query time, a metadata filter restricts retrieval to matching chunks, so one knowledge base can serve several product versions or departments. Only the sidecar file is read as knowledge base metadata; other attributes attached to the S3 object are not.
Parsing
The default parser extracts text. When PDFs contain tables, charts or images whose content matters, configure the data source with an advanced parser: a foundation model parser or Amazon Bedrock Data Automation. Chunking cannot recover what parsing dropped, so content lost from tables, figures or scanned pages has to be fixed at the parsing step, not by changing chunk sizes.
How do you mask, redact and anonymize sensitive data?
You mask, redact or anonymize sensitive data by transforming the values themselves in the dataset that will be used for training, with a managed detector where possible. Tools that only find sensitive data, block it at runtime, or encrypt it at rest do not produce a PII-free training set.
| Tool | What it does |
|---|---|
| Amazon Comprehend | Detects PII entities in text; an asynchronous PII job in redaction mode writes redacted copies of S3 documents, leaving originals untouched |
| AWS Glue Detect Sensitive Data transform | Finds PII in columns or text inside a Glue job and can redact (replace) what it finds before the output is written |
| AWS Glue DataBrew | Recipe steps for masking, substitution, keyed cryptographic hashing and deterministic encryption of columns |
| Amazon Transcribe | Can redact PII in transcripts of audio |
| AWS KMS encryption | Protects data at rest; anyone authorised to decrypt still sees the real values |
| Amazon Macie | Discovers and reports sensitive data in S3; it does not change the data |
| Amazon Bedrock Guardrails | Designed to mask or block PII in model inputs and outputs at runtime; it is not a training-data redaction job |
Pseudonymization that still joins
When identifiers must be hidden but datasets still have to be joined on them, apply the same deterministic transformation everywhere so identical IDs map to identical tokens. Non-deterministic substitution breaks the join. An unkeyed hash is deterministic, but for structured, guessable IDs anyone can hash every possible value and compare. A keyed hash (DataBrew's cryptographic hash uses a secret held in AWS Secrets Manager) is deterministic and cannot be reversed by someone who has only the prepared data.
How do you prepare data for fine-tuning, continued pre-training and distillation?
Prepare data for Amazon Bedrock model customization according to what you have: labeled examples go to fine-tuning, unlabeled domain text goes to continued pre-training, and prompts without answers plus a larger model to imitate go to distillation. Training data is JSON Lines (one JSON object per line) in Amazon S3, and the exact format depends on the method and the model.
| Method | You have | Record format |
|---|---|---|
| Fine-tuning (non-conversational text model) | Input-output pairs written or approved by people | {"prompt": "...", "completion": "..."} |
| Fine-tuning (conversational models, for example Amazon Nova) | Single or multi-turn conversations | For Amazon Nova: schemaVersion "bedrock-conversation-2024", a system block, and a messages array of alternating user and assistant turns (other models use their own system + messages layout) |
| Continued pre-training | Large amounts of unlabeled domain text | {"input": "..."} per document or passage |
| Distillation | Representative prompts, no reviewed answers | Prompts only (or invocation logs); the teacher model generates the responses |
In the conversation format, the assistant turns are what the model learns to produce, so map the agent's replies to the assistant role and the customer's messages to the user role; the roles must be exactly user and assistant, alternating, with the instruction in the system block. Continued pre-training teaches vocabulary and style, not a task, so it typically comes before task-specific fine-tuning. In distillation, the larger, stronger model is the teacher and the smaller, cheaper model is the student; Bedrock uses the teacher to create responses and fine-tunes the student on them, so you do not need to write answers yourself. Which models support each method varies by model and Region, so check the current list. Choosing among these strategies for a business need is covered in Task 2.1, and running and tuning the jobs in Task 2.2.
Tip. Task 1.2 questions are scenarios about a transformation pipeline, feature, stream, vector index or training dataset that has to meet a constraint: who builds it, how much infrastructure the team will run, how fresh or historically accurate the features must be, how late data can arrive, what a particular model can learn, how documents must be retrieved, which data must stay private, or what kind of training data exists. The deciding constraint is often stated in passing, so the work is to identify it and pick the service, setting or technique it implies, and to recognise when a plausible alternative changes the data, the scale or the meaning in a way the scenario does not allow. Some questions ask for more than one action that together complete a change.
- Glue = serverless Spark (job bookmarks for incremental runs); DataBrew = visual no-code recipes; EMR = cluster control and Spot task nodes; Data Wrangler = ML-focused visual prep you export to Pipelines, Feature Store or a processing job.
- Every feature group needs a record identifier and an event time; the online store serves the latest value in milliseconds, the offline store keeps every version in S3.
- Build training sets with point-in-time joins on the event time feature, never on current values or write_time; online-store TtlDuration expires stale records while the offline store keeps history.
- Lambda handles stateless per-record stream transforms and short tumbling windows; sliding or event-time windows with late data need Spark Structured Streaming with a watermark, or Flink.
- Scale features for distance and gradient-based models, use a robust scaler when outliers must stay, and fit every transform on the training split only, then reuse it at inference.
- Use log(1 + x) for right-skewed features with zeros, exponentiate predictions from a log target, and bin plus one-hot a non-monotonic feature for a linear model.
- One-hot encode nominal categories, ordinal-encode ordered ones, and split timestamps into local-time components.
- Documents and queries must share one embedding model and dimension; changing either means a new index and a full re-embed.
- Chunking: no chunking for short self-contained files, hierarchical for precision plus context, semantic threshold lower for smaller chunks, custom Lambda for your own splitting; metadata via <file>.metadata.json plus a sync.
- Redact the data itself (Comprehend, Glue sensitive data detection, DataBrew); Macie only discovers. Fine-tuning takes labeled pairs, continued pre-training unlabeled input text, distillation prompts and a teacher model.
Frequently asked questions
What is the difference between AWS Glue and AWS Glue DataBrew for ML data preparation?
AWS Glue runs code-based Spark ETL jobs without servers, suited to engineers who write PySpark or Scala and need features such as job bookmarks for incremental processing. AWS Glue DataBrew is a visual, no-code tool where analysts build cleaning and normalising steps as a recipe and rerun them as a scheduled recipe job. Both are serverless; the choice depends on whether the people building the transformation write code.
What is the difference between the SageMaker Feature Store online and offline stores?
The SageMaker Feature Store online store keeps only the latest version of each record and returns it by record identifier in milliseconds, for real-time inference. The offline store appends every version of every record to Amazon S3, for building training sets and keeping history. A feature group can use both, and the offline store is what makes point-in-time correct training data possible.
What is a point-in-time correct join in SageMaker Feature Store?
A point-in-time correct join matches each training label to the feature values that were current at the label's timestamp: the record version with the latest event time at or before that moment. It prevents feature leakage, where a model trains on information from after the event it predicts. In SageMaker Feature Store it runs against the offline store, keyed on the feature group's event time feature rather than the write time.
Should I use min-max normalization or standardization?
Standardization rescales a feature to mean 0 and standard deviation 1 and is the usual default for distance-based and gradient-based models. Min-max normalization rescales to a fixed range such as 0 to 1 and suits features without extreme outliers. If legitimate outliers must stay in the data, a robust scaler based on the median and interquartile range works better than either. Whichever you use, fit it on the training data only.
Which chunking strategy should I use for an Amazon Bedrock knowledge base?
Use default or fixed-size chunking for general documents, hierarchical chunking when you need small chunks for precise matching but larger parent chunks for context, semantic chunking when topics change unevenly within documents, and no chunking when each file is already a short self-contained answer. For splitting logic none of these can reproduce, choose no chunking and add a custom transformation Lambda function. The strategy is fixed once the data source is created.
Can I embed queries with a different model from the one used for my documents?
No. Vector search only works when documents and queries are embedded by the same model with the same settings. Different embedding models produce unrelated vector spaces even when their output dimensions match, so distances between a query vector and document vectors no longer reflect meaning. To switch models, re-embed every document with the new model into a new index and embed queries with that same model.
How do I remove PII from training data on AWS?
Transform the data itself before training: run an Amazon Comprehend asynchronous PII job in redaction mode to write redacted copies of text documents, use the AWS Glue Detect Sensitive Data transform to redact values inside a Glue job, or apply masking and keyed hashing steps in AWS Glue DataBrew. Amazon Macie only discovers sensitive data, and Amazon Bedrock Guardrails are designed for model inputs and outputs at runtime, so neither is a training-data redaction job.
What data does Amazon Bedrock need for fine-tuning, continued pre-training and distillation?
Fine-tuning needs labeled examples in JSON Lines: prompt and completion pairs, or, for conversational models, a system + messages format of alternating user and assistant turns (for Amazon Nova, schemaVersion bedrock-conversation-2024). Continued pre-training needs unlabeled text as input-only JSON Lines records. Distillation needs representative prompts; the larger teacher model generates the responses used to train the smaller student model.
Source
This lesson covers the "Data Preparation for ML and AI" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.
- AWS Certified Machine Learning Engineer – Associate (MLA-C02) exam guide — Amazon Web Services
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.