Collecting and Storing Data for ML and AI on AWS (MLA-C02)
Collecting and storing data for machine learning means getting raw data out of its sources, into the right AWS storage in the right format, and on to training jobs, feature stores and vector databases at the speed and cost the workload needs. On AWS that involves Amazon S3, EBS, EFS and FSx for Lustre, exports from Amazon RDS and DynamoDB, streaming with Amazon Kinesis, Amazon Data Firehose, Amazon MSK and Apache Flink, merging with AWS Glue and Apache Spark, vector stores in Amazon OpenSearch Service, pgvector and Amazon S3 Vectors, and Amazon SageMaker Feature Store. Task 1.1 of the MLA-C02 exam covers all of it: extracting data, choosing and securing storage, fixing capacity bottlenecks, ingesting streams, picking file formats, merging sources, configuring vector databases, storing text, images and audio, and ingesting features.
On this page10 sections
- How do SageMaker AI training jobs read data from Amazon S3, EFS and FSx for Lustre?
- How do you extract training data from RDS, Aurora and DynamoDB without loading production?
- How do you choose and configure storage by cost, performance, data structure and compliance?
- How do you troubleshoot ingestion and storage capacity problems?
- Which AWS streaming service should ingest ML data?
- Which data format fits each ML access pattern?
- How do you merge data from multiple sources with AWS Glue, Spark and code?
- How do you choose and configure a vector database for AI applications?
- How do you ingest and store text, image and audio data for ML and AI?
- How do you ingest data into SageMaker Feature Store and configure its stores?
- Configure SageMaker AI training input (File, FastFile, Pipe, EFS, FSx for Lustre) and data distribution for a given dataset and script
- Extract data from RDS, Aurora and DynamoDB into Amazon S3 without loading production
- Choose storage services and S3 storage classes by cost, performance and access pattern, and apply encryption, Object Lock, residency and Macie for compliance
- Diagnose hot shards, S3 throttling, small files and slow stream consumers, and pick the fix that addresses the cause
- Select between Data Firehose, Kinesis Data Streams, Amazon MSK and Managed Service for Apache Flink, and configure Firehose conversion, partitioning and buffering
- Choose Parquet, ORC, CSV, JSON Lines or Avro by access pattern and consumer, and merge sources efficiently with AWS Glue and Spark
- Select and configure a vector store, and ingest unstructured data and features into S3, Bedrock knowledge bases and SageMaker Feature Store
How do SageMaker AI training jobs read data from Amazon S3, EFS and FSx for Lustre?
A SageMaker AI training job reads each input channel either from Amazon S3, using one of three input modes, or from a file system data source on Amazon EFS or Amazon FSx for Lustre. The choice decides how long a job waits before its first epoch, how much local storage it needs, and whether the training script has to change.
| Channel setting | How data reaches the script | Good fit | Watch out for |
|---|---|---|---|
| S3, File mode (default) | The whole channel is downloaded to the instance's ML storage volume before the script starts; the script opens ordinary local files | Small and medium datasets, random access, scripts that expect local files | Start-up time grows with dataset size; the volume must hold the full download; every new job downloads again |
| S3, FastFile mode | Objects are exposed as a POSIX file system and streamed from S3 as the script reads them | Large datasets read mostly sequentially by code that already works with local file paths | Many tiny objects stay slow, because each open is an S3 request |
| S3, Pipe mode | Data is streamed through a Unix named pipe (FIFO); the script reads the pipe, not files | Legacy streaming readers written for pipes | Needs code that reads pipes; FastFile now covers most of its use cases with file semantics |
| File system data source: Amazon EFS | The job mounts an existing EFS file system | Data that already lives on EFS (for example shared with notebooks); no copy to S3 | The job needs VPC configuration with subnets and security groups that reach the mount targets |
| File system data source: FSx for Lustre | The job mounts a high-throughput parallel file system, typically linked to an S3 bucket through a data repository association | Throughput-bound, multi-instance, many-epoch training that repeats the same reads | Runs in a VPC; the file system lives in one Availability Zone, so the job's subnet should be in that zone |
Data distribution across instances
For S3 channels, the S3DataDistributionType decides what each instance of a multi-instance job receives. FullyReplicated (the default) gives every instance the entire channel, which suits model-parallel jobs or algorithms that need all data on each node, but multiplies download size and local storage. ShardedByS3Key splits the objects so each instance gets a different subset, which suits data-parallel jobs where each worker should see only its own slice of the data. Input mode and distribution are independent settings, so changing one never forces a change to the other.
Amazon EBS appears in this task mainly as the training instance's ML storage volume: in File mode the volume must be large enough for the dataset, and a larger volume does not remove the per-job download time.
How do you extract training data from RDS, Aurora and DynamoDB without loading production?
Use the managed export features, which read from backups or snapshots instead of the live database, so production serves no extra queries and consumes no extra capacity.
- DynamoDB export to Amazon S3 requires point-in-time recovery (PITR) to be enabled on the table. It reads from the continuous backup, so it consumes no read capacity and needs no code to page through items. Output is DynamoDB JSON or Amazon Ion in an S3 prefix. A full export captures the table at a point in time; an incremental export captures only items inserted, updated or deleted in a window you choose (from 15 minutes up to 24 hours), which is much cheaper than exporting a large table every day when only a small part changes. Exports have no built-in schedule; start them from a scheduler such as Amazon EventBridge Scheduler or a pipeline step.
- RDS and Aurora snapshot export to Amazon S3 extracts tables from an automated or manual DB snapshot (or an Aurora cluster) and writes them to S3 as Apache Parquet, ready for Athena, Glue or SageMaker. The running instance is not queried.
The general rule: any method that queries the live database, however it is scheduled, competes with production traffic. Change-capture features record only changes made after they are switched on and need consumer code, so on their own they cannot produce a complete historical copy. And a method whose output is another database, rather than files in S3, still leaves the export step undone.
Amazon OpenSearch Service is also a source: documents and logs indexed there can be read with the search APIs or snapshot tooling, but when it appears in Task 1.1 it is more often the destination, as a vector store (see below).
How do you choose and configure storage by cost, performance, data structure and compliance?
Pick the storage service by access pattern first, then the storage class by how often and how fast the data is read. Amazon S3 is the default home for training data of every type because it is durable, scales without limit and is readable in parallel by many jobs; file systems and block storage are chosen only when a workload needs their semantics or throughput.
| Need | Storage choice |
|---|---|
| Durable, cheap, shared store for datasets of any size and type | Amazon S3 |
| Shared POSIX files already used by notebooks or several jobs | Amazon EFS (its throughput mode must match the sustained load of the jobs reading it) |
| Maximum read throughput for repeated multi-instance training on S3 data | Amazon FSx for Lustre linked to the bucket |
| Single-digit-millisecond object access in one Availability Zone | S3 Express One Zone directory buckets |
| Per-instance scratch and File-mode downloads | The training instance's EBS ML storage volume |
S3 storage classes for ML data
| Class | Access | Cost profile | Fits |
|---|---|---|---|
| S3 Standard | Milliseconds | Highest storage price, no retrieval fee | Data in active use |
| S3 Intelligent-Tiering | Milliseconds in its default tiers (Frequent, Infrequent, Archive Instant Access) | Small per-object monitoring charge, no retrieval fees; moves each object automatically | Unknown or changing access patterns, such as experiment outputs whose future use nobody can predict |
| S3 Standard-IA / One Zone-IA | Milliseconds | Lower storage price, per-GB retrieval fee, minimum storage duration; One Zone-IA keeps one AZ copy | Known infrequent access |
| S3 Glacier Instant Retrieval | Milliseconds | Archive pricing with higher retrieval fees | Rarely read data that must still be instant |
| S3 Glacier Flexible Retrieval | Minutes to hours | Cheaper archive | Backups read a few times a year |
| S3 Glacier Deep Archive | Within about 12 hours (standard) or 48 hours (bulk) | Lowest storage price in S3 | Long-term retention, such as raw data held for years under a records-retention policy |
A lifecycle rule is the right tool when you know when data cools; Intelligent-Tiering is the right tool when you don't. Intelligent-Tiering's optional Archive Access and Deep Archive Access tiers must be activated separately, so its default configuration never reaches Deep Archive pricing.
Configuring storage for compliance
Compliance for training data comes down to four controls: who holds the encryption keys, whether data can be deleted or altered, where it physically lives and is processed, and whether you know which data is sensitive.
- Key control and audit. SSE-KMS with a customer managed key lets the company own the key policy, rotate and disable the key, and see every encrypt and decrypt call in AWS CloudTrail. SSE-S3 uses keys that Amazon S3 manages, so it does not meet a "we control the keys" requirement.
- Training job encryption. A training job stores data in two places, each with its own key setting:
KmsKeyIdinOutputDataConfigencrypts the model artifacts written to S3, andVolumeKmsKeyIdinResourceConfigencrypts the ML storage volumes attached to the instances. Instance types that use only NVMe instance storage encrypt it with hardware-managed keys and do not take a volume key. Keys configured on other resources, and in-transit encryption settings, do not encrypt a training job's own storage. - Immutability. S3 Object Lock (which requires versioning) keeps object versions from being deleted or overwritten for a retention period. Compliance mode cannot be shortened or bypassed by any user, including the root user. Governance mode can be bypassed by users granted
s3:BypassGovernanceRetention. Versioning alone keeps old versions but does not stop deletion. - Data residency. Data is processed where the job runs, and a SageMaker AI training job reads its dataset from the same Region as the job. To train in another Region you copy the data there first (for example with S3 Replication), and a residency rule limits which Regions are allowed for both the copy and the job. Residency is decided by Region choice alone; security controls protect data but do not change where it is stored or processed.
- Sensitive data discovery. Amazon Macie discovers personal and sensitive data in S3; automated sensitive data discovery samples objects continuously across buckets with managed data identifiers (personal identifiers, financial account details, credentials and more) and produces findings. Security services that watch access behaviour, resource configuration or software vulnerabilities do not look inside objects, so they cannot tell you where sensitive data lives.
How do you troubleshoot ingestion and storage capacity problems?
Most capacity failures in ML ingestion come from load concentrated on one unit of capacity (a shard, a prefix, a consumer) or from too many tiny requests, not from total capacity. Find the bottleneck from the symptom before adding resources.
| Symptom | Likely cause | Fix |
|---|---|---|
ProvisionedThroughputExceededException while aggregate stream traffic is modest and per-shard metrics show a single saturated shard | Low-cardinality partition key: Kinesis hashes each key to one shard, so a few keys use a few shards | Use a high-cardinality partition key (such as a device or session ID). Extra capacity does not help, because one key always hashes to one shard |
Rising IteratorAge where the Lambda consumer's processing time, not read throughput, is the constraint | Lambda processes one batch per shard at a time | Raise the event source mapping's ParallelizationFactor (1 to 10 concurrent batches per shard; order is kept per partition key). Match the fix to the constraint: a read-throughput fix does nothing for slow processing |
S3 503 Slow Down when a job writes millions of objects to one prefix | Request rate above the per-prefix limit (thousands of writes and reads per second per partitioned prefix) | Spread writes across more prefixes and write fewer, larger files. Retries with backoff smooth over throttling but do not raise the request-rate ceiling |
| Glue, Athena or training jobs slow to list and open data; many KB-sized objects | Small-file overhead: per-object request latency dominates | Compact output into larger files (roughly 100 MB or more), tune Firehose buffering, pack samples into shards |
| Training instances run out of disk in File mode | FullyReplicated copies everything to every instance | ShardedByS3Key, FastFile mode, or a file system source |
Kinesis capacity facts worth knowing: each shard accepts up to 1 MB per second or 1,000 records per second of writes and serves 2 MB per second of reads shared by standard consumers; on-demand mode scales shard count automatically but still hashes keys to shards.
Which AWS streaming service should ingest ML data?
Choose by what has to happen to the data in flight: Amazon Data Firehose when records only need delivering, Amazon Kinesis Data Streams when applications must read and replay them, Amazon MSK when producers speak Apache Kafka, and Amazon Managed Service for Apache Flink when the stream must be processed with state and time windows.
| Service | What it is | Choose it when | It cannot |
|---|---|---|---|
| Amazon Data Firehose | Fully managed delivery: buffers records and writes them to S3, Redshift, OpenSearch, Apache Iceberg tables and other destinations; optional Lambda transform and format conversion | Data only has to land somewhere within minutes, with no consumer code to run | Be read by your own applications, replay data, or compute aggregates |
| Amazon Kinesis Data Streams | Durable, ordered stream of shards; records retained for 24 hours by default, extendable up to 365 days | Several applications read the same records, need replay after a bug fix, or need low latency | Deliver to S3 by itself (it needs a consumer or Firehose) |
| Amazon MSK | Managed Apache Kafka (provisioned or Serverless), with MSK Connect for connectors | Producers or consumers already use Kafka APIs and cannot change | Accept Kinesis API producers |
| Amazon Managed Service for Apache Flink | Managed Apache Flink applications reading from Kinesis or Kafka | Windowed aggregates, joins, durable exactly-once state via checkpoints | Act as the ingestion buffer itself |
Consumers on Kinesis Data Streams
Standard (shared-throughput) consumers poll with GetRecords and split each shard's 2 MB per second of read throughput between them, with latency around 200 ms or more per consumer. Enhanced fan-out registers each consumer for its own dedicated 2 MB per second per shard, pushed over HTTP/2 with latency around 70 ms. Enhanced fan-out solves a throughput and latency problem between consumers; the retention period solves a replay problem. They are separate settings, so a requirement for both needs both. A message queue deletes messages once they are consumed, so it cannot stand in for a stream that several readers replay.
Configuring Firehose delivery for ML datasets
Firehose shapes what lands in S3 through four settings: record format conversion, data transformation with Lambda, dynamic partitioning, and buffering hints.
- Record format conversion writes Apache Parquet or ORC instead of the raw records, using a table in the AWS Glue Data Catalog as the schema. It accepts JSON input only (deserialized with the OpenX JSON SerDe or Hive JSON SerDe). CSV or other formats must first be turned into JSON by a Lambda transformation. A transformation function returns records, not files: Firehose concatenates whatever it returns into its own objects, so producing columnar files is the conversion feature's job.
- Data transformation invokes a Lambda function on each buffered batch; the function returns every record with its
recordId, aresult(Ok, Dropped or ProcessingFailed) and base64-encoded data. It changes record content, not delivery timing. - Dynamic partitioning groups objects under S3 prefixes built from values inside each record, extracted with inline JSON parsing (jq expressions) or a Lambda function and referenced in the prefix with
!{partitionKeyFromQuery:...}. Without dynamic partitioning, a prefix cannot use values from inside the record: its namespaces are!{timestamp:...}, which groups objects by arrival time, and!{firehose:...}(a random string, or the error type in an error output prefix). Partitioning on a high-cardinality key multiplies the number of objects. - Buffering hints set a size (up to 128 MiB) and an interval (up to 900 seconds). Firehose writes an object when either limit is reached first. To make fewer, larger objects, raise the limit that is currently triggering the flush, then check that the other limit does not simply take over.
A worked example: a stream carrying about 300 MB per minute with hints of 64 MiB and 300 seconds reaches the size limit every 13 seconds or so, so the interval never matters. Lengthening the interval changes nothing; only a larger size hint (up to 128 MiB) produces bigger objects, and at that rate the interval still never fires.
Which data format fits each ML access pattern?
Use a columnar format (Apache Parquet or ORC) for data that is written once and read by column for analytics and training, and a row-based format (CSV, JSON Lines, Avro) for data written or exchanged one record at a time or required by a specific consumer.
| Format | Layout | Strengths | Typical ML use |
|---|---|---|---|
| Apache Parquet | Columnar, binary, compressed (Snappy is common), with per-column statistics | Reads only the requested columns and skips row groups, so Athena (billed per byte scanned) and Spark read far less | Curated training layers, feature tables, exports; often partitioned by date in S3 |
| ORC | Columnar, binary | Similar column pruning and predicate pushdown; common in Hive ecosystems | Firehose conversion target, Hive and Spark tables |
| CSV | Row text | Universal and simple; no schema or types | Inputs for SageMaker AI built-in algorithms |
| JSON / JSON Lines | Row text, one object per line for JSON Lines | Nested, self-describing records | Event payloads; Amazon Bedrock fine-tuning datasets |
| Apache Avro | Row binary with a schema | Compact, fast to write per record, schema evolution | Streaming messages in Kafka or Kinesis |
Format rules that trip people up
- SageMaker AI built-in algorithms with
text/csv(for example XGBoost and Linear Learner) expect the target in the first column and no header row. A file with the target last trains a model that predicts whichever feature sits in the first column. Many built-in algorithms also accept RecordIO-protobuf, and XGBoost accepts LIBSVM and Parquet. - Amazon Bedrock model customization takes training and validation data from S3 as JSON Lines: one JSON object per line, such as prompt and completion fields (or a messages structure for conversational models). The format is line-delimited, so every line must parse as a complete JSON object on its own.
- Schema enforcement on streams. AWS Glue Schema Registry stores schemas (Avro, JSON Schema or Protobuf) and applies compatibility modes such as BACKWARD, rejecting a producer's new schema version that would break consumers. Avro and Protobuf also give compact binary messages, while JSON Schema validates structure but keeps text payloads. Data without a registered schema gets no compatibility checks.
Pipelines usually combine both families: a row format where records are produced, and a columnar format, partitioned by the column most queries filter on, where they are analysed.
How do you merge data from multiple sources with AWS Glue, Spark and code?
Merge large sources with a distributed Apache Spark engine, and pick where Spark runs by how much infrastructure the team will manage: AWS Glue runs serverless Spark jobs that read Data Catalog tables directly, including JDBC sources such as Amazon RDS, while Amazon EMR runs Spark on clusters you configure. Single-node tools are fine for small data but do not scale to terabytes.
Making a Glue join fast
- Prune partitions before reading. A
push_down_predicateon a partition column (for example the date) filters partitions in the Data Catalog, so Glue never lists or reads the others, and it works for any date you rerun. - Parallelise JDBC reads. Glue reads a JDBC table through one connection by default, which shows up as a single task. Setting
hashfield(orhashexpression) withhashpartitionssplits the read into parallel queries. Because the bottleneck is one connection, capacity added elsewhere in the job stays idle. - Job bookmarks process only data added since the last run. They suit incremental pipelines; reprocessing a chosen historical date is a job for partition predicates.
Join strategy in Spark
Spark's default sort-merge join shuffles both tables by the join key, which is expensive when one side is terabytes. When the other side is small enough to fit in executor memory, a broadcast hash join sends the small table to every executor and avoids shuffling the large one. Spark broadcasts automatically only below spark.sql.autoBroadcastJoinThreshold (10 MB by default), so a larger reference table needs a broadcast hint or a higher threshold. Settings that tune the shuffle make it cheaper but still move the large table across the network; only a broadcast avoids that.
Amazon Athena can also join sources, including federated queries to databases through a Lambda-based connector, but the connector is something to deploy and a CTAS or INSERT INTO query can write at most 100 partitions, so Athena suits ad hoc merges more than large nightly builds.
How do you choose and configure a vector database for AI applications?
Choose the vector store by query rate and latency, by whether vectors must sit beside relational data, by the retrieval features needed (such as hybrid keyword plus vector search), and by cost; then configure the index so its dimension, distance metric and algorithm match the embedding model and the workload.
| Store | Strengths | Best fit | Limits |
|---|---|---|---|
| Amazon S3 Vectors (vector buckets and vector indexes) | No infrastructure; pay for storage and queries; metadata filtering; sub-second latency for infrequent queries and as low as about 100 ms for more frequent ones; integrates with Amazon Bedrock Knowledge Bases | Very large, cost-sensitive collections with infrequent queries where sub-second latency is acceptable | Similarity search only (no keyword or hybrid search of its own); AWS positions it for workloads where queries are less frequent. OpenSearch Service can use S3 Vectors as a storage engine when hybrid search is needed |
| Amazon OpenSearch Serverless, vector search collection | Serverless k-NN with millisecond latency at high query rates; combines lexical and vector queries (hybrid search) | Production RAG or semantic search with high traffic and no clusters to manage | Higher baseline cost than S3 Vectors; a search collection type does not support vector fields |
| Amazon OpenSearch Service managed domain (k-NN) | Full control of instances, engines and plugins | Teams that already run OpenSearch or need custom tuning | You size, patch and scale the domain |
| Amazon Aurora PostgreSQL or RDS for PostgreSQL with pgvector | Vectors in the same tables as business data; one SQL query combines similarity with relational filters on committed data | Catalogs, inventories and apps already on PostgreSQL whose similarity results must reflect committed changes immediately | Scaling is database scaling; not a search engine |
Any store that holds a copy of operational data (OpenSearch, S3 Vectors) must be kept in sync, so its results lag committed changes; pgvector avoids that when the source is PostgreSQL. Amazon Bedrock Knowledge Bases also support other stores, such as Amazon Neptune Analytics and several partner databases.
Configuring the index to specification
- Dimension is fixed when the vector field or index is created and must equal the embedding model's output. Amazon Titan Text Embeddings V2 returns 1,024 dimensions by default (configurable to 512 or 256); the older Titan Embeddings G1 - Text returns 1,536. A mismatch makes every write fail; the fix is a new field or index (or a model whose output matches), because no other index setting changes its dimension.
- Distance metric (cosine, Euclidean/L2, inner product) should match how the model was trained; S3 Vectors supports cosine and Euclidean and fixes the metric at index creation.
- Algorithm. HNSW builds a graph that gives a strong speed-to-recall trade-off with no training step and can be built on an empty table; IVFFlat clusters vectors into lists, builds faster and uses less memory but needs representative data present before the index is built and usually recalls less.
Prerequisites for Amazon Bedrock knowledge base vector stores
- OpenSearch (Serverless or managed domain): create the index with a
knn_vectorfield whose dimension matches the model and whose engine is faiss (the documentation states nmslib is not supported for managed domains), a text field for the chunk text, and the Bedrock-managed metadata field. Source details such as the S3 URI are kept in that metadata field. - Aurora PostgreSQL with pgvector: the cluster must be in the same account as the knowledge base, with the RDS Data API enabled; create a table with an id (primary key), an embedding vector column, a chunks text column, a metadata column and optionally a custom_metadata column; create an HNSW index on the embedding with
vector_cosine_ops, a GIN index on the chunks text (for hybrid search) and on custom_metadata if it is used for filtering; and store the database credentials in an AWS Secrets Manager secret.
How do you ingest and store text, image and audio data for ML and AI?
Store unstructured objects (documents, images, audio, video) in Amazon S3 and keep their searchable attributes in a database or alongside the objects, with the S3 key as the link between them. S3 handles large objects, unlimited growth and highly parallel reads; databases are poor object stores (a DynamoDB item is limited to 400 KB, and relational BLOB columns are costly and slow at scale), and a single EBS volume is neither shared nor elastic.
- Metadata. Put per-object attributes that are queried by key (for example capture date, device, class label or owning team) in DynamoDB or a relational table, holding the S3 key. Manifest files list objects and their labels for a training channel or labeling job without moving data.
- Many small files. Millions of KB-sized images make training spend its time opening objects, even in FastFile mode. Pack samples into shard files of a few hundred MB, such as tar archives in WebDataset style, RecordIO or TFRecord, so loaders read sequentially. The fix has to reduce the number of objects; changes that leave the object count the same do not help.
- Documents for Amazon Bedrock Knowledge Bases. For an S3 data source, custom metadata comes from a sidecar file stored beside each document and named after it with the
.metadata.jsonsuffix (for examplepolicy.pdf.metadata.json, holding ametadataAttributesobject). Attributes are indexed with the chunks during sync and can be used as metadata filters at query time. For this data source the sidecar file is where ingestion reads custom metadata from; attributes kept anywhere else are not picked up. Chunking and parsing choices belong to Task 1.2. - Audio and images feeding AI services. Amazon Transcribe, Amazon Textract and Amazon Rekognition read their inputs from S3 and can write outputs back to S3, so the same bucket layout serves both training and AI-service pipelines.
How do you ingest data into SageMaker Feature Store and configure its stores?
Ingest into a SageMaker Feature Store feature group with the PutRecord API for streaming, low-latency updates, and with batch ingestion (the Feature Store Spark connector, the SDK's FeatureGroup.ingest(), or a SageMaker Data Wrangler export) for bulk loads. Where the data lands depends on the feature group's online and offline store configuration.
| Online store | Offline store | |
|---|---|---|
| Holds | Only the latest record per record identifier (the one with the newest event time) | Every record ever ingested, append-only |
| Read with | GetRecord / BatchGetRecord in milliseconds (Standard tier; an InMemory tier offers lower latency) | SQL with Athena or Spark over Parquet files in S3, registered in the Glue Data Catalog |
| Used for | Real-time inference lookups | Building training sets with full history |
| Freshness | Immediate after PutRecord | PutRecord data is buffered and written to S3 within about 15 minutes |
A feature group can have either store or both; latest-value lookups and full history need both. Every record must include a record identifier feature and an event time feature. Because the online store keeps the newest event time, writing an older event time (for example when a backfill or a delayed producer writes an earlier timestamp) succeeds but does not overwrite a newer online value; the offline store keeps both records.
Choosing the ingestion path
- Streaming features (for example a Lambda function consuming Kinesis): call PutRecord per updated entity. The online store is updated at once and, with an offline store configured, the record is also replicated to S3. Records enter a feature group through the Feature Store APIs or connectors; files placed in the offline store's S3 location by other means are not feature group records.
- Backfills of billions of rows: run the Spark connector on Amazon EMR (or Glue/Processing with Spark) and set the target store. Targeting
OfflineStoreloads history for training without pushing stale values into the online store; single-process ingestion is meant for modest volumes, not billions of rows. - Read-after-write in batch pipelines: offline replication of PutRecord data takes up to about 15 minutes, so SQL over the offline store reflects writes only after that delay. A pipeline that reads its own recent writes from the offline store has to allow for it (or read the latest values from the online store).
Store configuration that is fixed or easy to miss
- Offline table format is AWS Glue (default) or Apache Iceberg, chosen only at creation. With Iceberg, the many small Parquet files that frequent PutRecord calls produce can be compacted with Athena
OPTIMIZE ... REWRITE DATA. To move an existing Glue-format feature group to Iceberg, create a new feature group and backfill it. Offline store partitioning is managed by Feature Store, so compaction, not re-partitioning, is the lever for small files. - Time to live (TTL) applies to the online store only. Set
TtlDurationas a feature-group default in the online store configuration, or per record in a PutRecord call; a record expires at its event time plus the duration. Expired records are hard-deleted from the online store, and the deletion is recorded in the offline store, so training history stays. A feature-group default applies to every record with no change to the code that writes them, and expiry needs no scheduled clean-up job. Training history survives because TTL never deletes from the offline store.
Building point-in-time correct training sets from the offline store, and managing feature definitions, are covered in the Task 1.2 lesson on feature engineering.
Tip. Task 1.1 is tested with short scenarios about ML data pipelines: a training job that starts slowly or runs short of disk, an extraction that must not burden production, a stream or bucket that throttles or falls behind, data that must be cheaper to keep or must meet a compliance rule, a vector store that has to match an embedding model, unstructured data that must be ingested and made searchable, or features that must be served online and kept for training. Each scenario states one or two constraints, about code changes, operational effort, latency, cost, data location or freshness, and the right answer is the option that meets every stated constraint rather than the most powerful service. Options often differ in a single configuration detail, so knowing what each setting actually does, and which settings are fixed at creation, matters more than recognising service names. Some questions ask for more than one action.
- File mode downloads the whole channel before training; FastFile streams with file semantics and no script change; Pipe needs pipe-reading code. ShardedByS3Key splits data across instances, FullyReplicated copies all of it to each.
- EFS and FSx for Lustre channels need VPC configuration; FSx for Lustre is the throughput choice for repeated multi-instance training and lives in one Availability Zone.
- DynamoDB export to S3 needs PITR and uses no read capacity (full or incremental); RDS and Aurora snapshot export writes Parquet without querying the instance.
- Intelligent-Tiering for unknown access with no retrieval fees; Glacier Deep Archive for the cheapest long-term retention; Object Lock compliance mode for immutability nobody can bypass; SSE-KMS with a customer managed key for key control.
- A training job reads data from its own Region and has separate key settings for output artifacts (KmsKeyId) and instance volumes (VolumeKmsKeyId).
- Hot Kinesis shards need a better partition key, not more shards; slow Lambda consumers need ParallelizationFactor; S3 503s need more prefixes and fewer, larger files.
- Firehose for delivery without code, Kinesis Data Streams for replay and multiple consumers (enhanced fan-out for dedicated throughput), MSK for Kafka producers, Managed Flink for stateful windows.
- Firehose format conversion needs JSON input and a Glue table; dynamic partitioning groups by record values; an object is written when either buffering hint is reached.
- Parquet or ORC for column reads; built-in algorithm CSV has the target first and no header; Bedrock fine-tuning data is JSON Lines; Avro plus Glue Schema Registry enforces stream schema compatibility.
- S3 Vectors for cheap infrequent queries, OpenSearch Serverless vector search for high-rate hybrid search, pgvector for vectors next to live relational data; index dimension must equal the embedding model's output.
Frequently asked questions
What is the difference between File, FastFile and Pipe mode in SageMaker AI training?
File mode, the default, downloads the whole S3 channel to the training instance's storage volume before the script starts. FastFile mode presents the S3 objects as local files but streams them on demand, so training starts almost immediately without changing a script that opens file paths. Pipe mode streams data through a named pipe and requires code that reads from the pipe; FastFile now covers most cases where Pipe was used.
How do you copy a DynamoDB table to Amazon S3 for training without using read capacity?
Enable point-in-time recovery on the table and use DynamoDB export to S3. The export reads from the continuous backup rather than the live table, so it consumes no read capacity and needs no scanning code. Use a full export for a complete copy, or an incremental export of a 15-minute to 24-hour window when only recent changes are needed. Exports are started on demand, so a scheduler triggers recurring ones.
When should I use Amazon Data Firehose instead of Kinesis Data Streams for ML data?
Use Amazon Data Firehose when records only need to be delivered to a destination such as Amazon S3 within minutes, optionally converted to Parquet or transformed by Lambda, with no consumer code to run. Use Kinesis Data Streams when your own applications must read the records, several consumers need the same data, records must be replayed after a fix, or latency must be well under a minute.
Why does Firehose record format conversion fail on CSV data?
Firehose record format conversion to Parquet or ORC accepts only JSON input, which it deserializes with a JSON SerDe and maps to a schema from an AWS Glue Data Catalog table. CSV records must first be converted to JSON by a Firehose Lambda data transformation; the conversion step then writes columnar output.
Which vector database should I use for a RAG application on AWS?
Use Amazon S3 Vectors for very large, cost-sensitive collections queried infrequently where sub-second latency is fine. Use an Amazon OpenSearch Serverless vector search collection for high query rates, millisecond latency or hybrid keyword and vector search without managing clusters. Use Aurora or RDS for PostgreSQL with pgvector when the vectors must be queried together with live relational data in SQL. In every case the index dimension must equal the embedding model's output size.
Why do Amazon Bedrock knowledge base syncs fail with a vector dimension error?
The vector field or index was created with a different dimension from the embedding model's output. Amazon Titan Text Embeddings V2 returns 1,024 dimensions by default (or 512 or 256 if configured), while the older Titan Embeddings G1 - Text returns 1,536. Dimensions are fixed at creation, so the field must be recreated to match, or a model with matching output used.
What is the difference between the online and offline store in SageMaker Feature Store?
The online store keeps only the latest record for each record identifier and serves it in milliseconds through GetRecord for real-time inference. The offline store appends every record to Amazon S3 as Parquet with a Glue Data Catalog table, so the full history can be queried with SQL for training. Records written with PutRecord reach the online store immediately and the offline store within about 15 minutes.
How do you load years of historical features into SageMaker Feature Store?
Use batch ingestion with the Feature Store Spark connector, for example on Amazon EMR, which distributes the writes across a cluster. If only training needs the history, set the target store to the offline store so the online store keeps holding just the latest values from live ingestion. Single-process ingestion is meant for modest volumes, and records must enter through the Feature Store APIs or connectors.
Source
This lesson covers the "Data Preparation for ML and AI" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.
- AWS Certified Machine Learning Engineer – Associate (MLA-C02) exam guide — Amazon Web Services
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.