SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
MLA-C02 quick-recall

MLA-C02 flashcards

Flip through 12 cards — one per MLA-C02 topic — and self-test the key exam facts. Free, no account needed. These exams reward fast recognition, which is exactly what flashcards train.

1 / 12

Every MLA-C02 flashcard, by exam domain

120 key facts across 4 domains — the full deck below, so you can scan it even without the interactive cards.

Data Preparation for ML and AI

28% of the exam
  • File mode downloads the whole channel before training; FastFile streams with file semantics and no script change; Pipe needs pipe-reading code. ShardedByS3Key splits data across instances, FullyReplicated copies all of it to each.
  • EFS and FSx for Lustre channels need VPC configuration; FSx for Lustre is the throughput choice for repeated multi-instance training and lives in one Availability Zone.
  • DynamoDB export to S3 needs PITR and uses no read capacity (full or incremental); RDS and Aurora snapshot export writes Parquet without querying the instance.
  • Intelligent-Tiering for unknown access with no retrieval fees; Glacier Deep Archive for the cheapest long-term retention; Object Lock compliance mode for immutability nobody can bypass; SSE-KMS with a customer managed key for key control.
  • A training job reads data from its own Region and has separate key settings for output artifacts (KmsKeyId) and instance volumes (VolumeKmsKeyId).
  • Hot Kinesis shards need a better partition key, not more shards; slow Lambda consumers need ParallelizationFactor; S3 503s need more prefixes and fewer, larger files.
  • Firehose for delivery without code, Kinesis Data Streams for replay and multiple consumers (enhanced fan-out for dedicated throughput), MSK for Kafka producers, Managed Flink for stateful windows.
  • Firehose format conversion needs JSON input and a Glue table; dynamic partitioning groups by record values; an object is written when either buffering hint is reached.
  • Parquet or ORC for column reads; built-in algorithm CSV has the target first and no header; Bedrock fine-tuning data is JSON Lines; Avro plus Glue Schema Registry enforces stream schema compatibility.
  • S3 Vectors for cheap infrequent queries, OpenSearch Serverless vector search for high-rate hybrid search, pgvector for vectors next to live relational data; index dimension must equal the embedding model's output.
  • Glue = serverless Spark (job bookmarks for incremental runs); DataBrew = visual no-code recipes; EMR = cluster control and Spot task nodes; Data Wrangler = ML-focused visual prep you export to Pipelines, Feature Store or a processing job.
  • Every feature group needs a record identifier and an event time; the online store serves the latest value in milliseconds, the offline store keeps every version in S3.
  • Build training sets with point-in-time joins on the event time feature, never on current values or write_time; online-store TtlDuration expires stale records while the offline store keeps history.
  • Lambda handles stateless per-record stream transforms and short tumbling windows; sliding or event-time windows with late data need Spark Structured Streaming with a watermark, or Flink.
  • Scale features for distance and gradient-based models, use a robust scaler when outliers must stay, and fit every transform on the training split only, then reuse it at inference.
  • Use log(1 + x) for right-skewed features with zeros, exponentiate predictions from a log target, and bin plus one-hot a non-monotonic feature for a linear model.
  • One-hot encode nominal categories, ordinal-encode ordered ones, and split timestamps into local-time components.
  • Documents and queries must share one embedding model and dimension; changing either means a new index and a full re-embed.
  • Chunking: no chunking for short self-contained files, hierarchical for precision plus context, semantic threshold lower for smaller chunks, custom Lambda for your own splitting; metadata via <file>.metadata.json plus a sync.
  • Redact the data itself (Comprehend, Glue sensitive data detection, DataBrew); Macie only discovers. Fine-tuning takes labeled pairs, continued pre-training unlabeled input text, distillation prompts and a teacher model.
  • Glue Data Quality in an ETL job gives row-level results so bad rows can be quarantined in the same run; Data Catalog runs check data at rest, recommend rules and filter with partition predicates.
  • DQDL Boolean rules (IsComplete, IsUnique) allow zero failures; ratio rules (Completeness, Uniqueness) take a threshold; last(k) with k > 1 needs an aggregation such as min or avg.
  • A DataBrew ruleset is evaluated by a profile job, and the pass or fail lives in the Ruleset Validation Result event, not the job state.
  • SageMaker Ground Truth and SageMaker Clarify are closed to new customers; a new account labels with FM pre-labeling plus human audit and computes bias metrics itself.
  • Split by key for repeated entities, use an ordered split for time series, shuffle sorted exports and deduplicate before any split.
  • CI measures how many rows each facet has; DPL measures how often each facet gets the positive label; CDD conditions on a subgroup to expose Simpson's paradox.
  • Resample only the training split after splitting; when rows must stay as recorded, use scale_pos_weight (binary), XGBoost instance weights with csv_weights=1 (multi-class) or inverse-frequency loss weights.
  • Augment rare conditions with label-preserving transforms in training only, and measure on real held-out examples.
  • ApplyGuardrail screens text against an existing guardrail without invoking a model; Comprehend gives per-category toxicity scores; Rekognition DetectModerationLabels screens images.
  • Impute skewed numbers with the median and categories with the mode, fitted on the training split; treat sentinel codes as missing and keep informative extremes.

ML Model and Foundation Model (FM) Development

24% of the exam
  • Use the least custom approach that works: AI service, then FM or traditional ML, then a custom model; FMs are a poor fit for high-volume tabular or fixed-class prediction when labels exist.
  • Interpretability requirements decide the model: fixed published weights mean a linear model; SHAP values vary per prediction, importance rankings lack direction, monotonic constraints fix direction only.
  • Filter FMs on hard requirements (modality, context window, languages, Region) first, then compare candidates on your own prompts for quality, latency and cost.
  • The same embedding model must embed documents and queries; use a multimodal model on the images for visual search and a multilingual model for cross-language retrieval; fewer dimensions means less storage.
  • RAG handles changing facts, citations and access control; fine-tuning handles tone, layout and labeling behaviour; many systems need both.
  • Customization follows the data: labeled pairs for supervised fine-tuning, a reward function for reinforcement fine-tuning, a good teacher and prompts for distillation, unlabeled text for continued pre-training.
  • LoRA adapters cut training cost and let many variants share one base model on a SageMaker AI endpoint.
  • Hybrid search fixes exact identifiers, GraphRAG fixes multi-hop questions, structured data stores compute exact aggregates; a missing mandatory feature rules out a cheaper vector store.
  • Prompt caching, batch inference, Provisioned Throughput, prompt routing and cross-Region inference each solve one cost or capacity problem and carry one constraint.
  • Textract AnalyzeExpense for invoices, AnalyzeID for identity documents, Rekognition DetectModerationLabels for unsafe images, Transcribe for speech with PII redaction and custom vocabulary, Comprehend for text.
  • Built-in algorithm choice follows labels and output: RCF for unlabeled anomaly scores, DeepAR for many related time series, XGBoost for tabular prediction, k-means only clusters.
  • Built-in CSV input means the label in the first column and no header; FastFile streams data and ShardedByS3Key gives each instance its own subset.
  • In script mode, hyperparameters arrive as command-line arguments, the model must be saved to SM_MODEL_DIR (/opt/ml/model), and requirements.txt in source_dir adds packages to the managed container.
  • Bayesian tuning learns from completed jobs, so high parallelism turns it into random search; use logarithmic scaling for ranges spanning orders of magnitude and a regex metric definition for custom scripts.
  • Warm start TransferLearning handles new data or a new algorithm version; IdenticalDataAndAlgorithm needs the same data and image.
  • Early stopping (AMT Auto, Hyperband, XGBoost early_stopping_rounds with a validation channel) saves wasted compute; use data parallelism when the model fits on one GPU and model parallelism or sharding when it does not.
  • Overfitting needs regularization or data, underfitting needs capacity or features, and catastrophic forgetting needs LoRA, mixed general data, a lower learning rate or fewer epochs.
  • Bagging reduces variance, boosting reduces bias, stacking learns how to weight diverse models, and cascades or Bedrock intelligent prompt routing cut cost by sending easy requests to smaller models.
  • Prompt engineering first; fine-tuning needs labeled pairs, continued pre-training uses unlabeled text, and distillation trains a smaller student from a teacher's responses to your prompts.
  • Changing the embedding model or dimension means re-embedding the corpus; hybrid search fixes exact identifiers, hierarchical chunking adds context, and overlap stops split sentences.
  • Reproducible runs need the data version (MLflow dataset input with a digest) and the code version (Git commit tag), not only parameters and metrics; a training job reaches managed MLflow only with the tracking server ARN as tracking URI and the sagemaker-mlflow plugin.
  • In Bedrock Prompt Management, iterate on the draft, compare alternatives as variants, and ship immutable numbered versions by ARN; fix a prompt by creating a new version, never by editing or deleting the old one.
  • Baselines come from the training data. Input-distribution checks catch data drift; only ground-truth quality metrics, joined by inference ID when labels arrive late, catch concept drift.
  • SageMaker Model Monitor, Clarify and Debugger are closed to new customers: new accounts use data capture with Evidently, MLflow and SNS/CloudWatch for drift, the SHAP library and pandas or scikit-learn bias formulas for predictive models, fmeval or Bedrock evaluations for foundation models, and metric definitions, CloudWatch alarms and TensorBoard for training.
  • A shadow variant sees live traffic but never answers customers; it needs an instance-based real-time endpoint, data capture compares responses, and completing the test can deploy the shadow variant.
  • Local explanation means one prediction's SHAP values; global ranking means mean absolute SHAP; how the prediction changes as one feature moves is a partial dependence plot.
  • Divergence to NaN: lower the learning rate. Vanishing gradients: ReLU with He initialization. Exploding gradients: clip.
  • With rare positives, ignore accuracy and compare recall, precision or PR AUC; linear error cost means MAE, costly large misses mean RMSE.
  • BLEU for translation, ROUGE for summaries, BERTScore or embedding similarity when correct answers are paraphrased.
  • Bedrock evaluations: automatic for repeatable algorithmic scores, LLM-as-a-judge (with custom metrics and an independent, human-calibrated judge) for subjective quality, human jobs with your own team for expert or confidential review, retrieve-only RAG jobs to test retrieval and retrieve-and-generate jobs for faithfulness and citations.

Deployment and Orchestration of ML and AI Workflows

24% of the exam
  • Synchronous plus idle gaps plus acceptable cold start means serverless inference; large payloads or minutes of processing with scale-to-zero means asynchronous inference; a scheduled dataset means batch transform.
  • Asynchronous requests pass the payload as an S3 InputLocation, and success and error SNS topics in AsyncInferenceConfig report completion without polling.
  • Batch transform needs SplitType Line with MultiRecord batching for big files, parallelizes by file, and filters with InputFilter, JoinSource and OutputFilter in that order.
  • Inferentia2 is the low-cost inference accelerator only for Neuron-compatible models, Graviton needs an arm64 image, and GPU instances should be right-sized to the model's memory.
  • Multi-model endpoints share one container across many similar models (Triton on GPUs); multi-container Direct mode hosts different frameworks; serial pipelines chain containers; inference components give each model its own resources and scaling.
  • Large models fit only when weights plus KV cache fit the usable GPU memory: shard with tensor parallelism across every GPU, or quantize when the instance is fixed.
  • Bedrock on-demand bills per token, cross-Region inference profiles absorb peaks without commitment, Provisioned Throughput is a commitment, batch inference serves offline jobs, and Marketplace models run on endpoints you size.
  • Custom Model Import takes supported LLM architectures in Hugging Face safetensors format and serves them on demand, but not embedding models and not with batch inference.
  • Bedrock Agents act through action groups (Lambda or return of control, with optional user confirmation); AgentCore Runtime hosts framework agents; MCP connects agents to tools and A2A connects agents to agents.
  • Reranking fixes good chunks ranked too low, metadata filters fix wrong-scope answers, and query decomposition fixes one-sided multi-part answers.
  • Only Bedrock Provisioned Throughput reserves model capacity for one workload; cross-Region inference profiles absorb bursts while staying on-demand, and batch inference is never interactive.
  • Serverless provisioned concurrency removes cold starts and can be scheduled with Application Auto Scaling, but its scalable target floors at 1, and serverless has no GPUs.
  • CloudFormation exports lock the producer while imported; Parameter Store dynamic references do not. In Step Functions, SageMaker AI .sync exists only for jobs, never .waitForTaskToken, so wait for endpoints with a Wait + DescribeEndpoint + Choice loop.
  • Inference containers answer GET /ping and POST /invocations on port 8080; training scripts must save the model to /opt/ml/model. Extend the AWS image FROM it when you need OS packages or have no internet.
  • VpcConfig belongs on the model. Only S3 and DynamoDB have gateway endpoints; runtime and control-plane APIs (sagemaker.runtime vs sagemaker.api, bedrock-runtime vs bedrock) are separate interface endpoints, and default SDK hostnames need private DNS.
  • Deploy with create_model, create_endpoint_config, create_endpoint; configurations are immutable, so swap versions with a new configuration and update_endpoint, and wait with the endpoint_in_service waiter on the endpoint name.
  • Scale short requests on invocations per instance, streaming LLMs on high-resolution concurrent requests, async on backlog, and variable-cost GPU work on GPUUtilization from /aws/sagemaker/Endpoints; scale from zero needs a step policy on NoCapacityInvocationFailures (or HasBacklogWithoutCapacity for async).
  • Inference components plus managed instance scaling give each model its own GPUs and scaling; training plans reserve GPU capacity, which quotas, Spot and Savings Plans do not.
  • Knowledge base vector fields must match the embedding model's dimensions; hybrid search needs OpenSearch Serverless, RDS/Aurora or MongoDB; Aurora needs the Data API and a Secrets Manager secret; metadata comes from .metadata.json sidecars.
  • AgentCore Runtime sessions are ephemeral ARM64 microVMs; replay exact turns from Memory events, keep distilled preferences with long-term strategies, and use Gateway for shared MCP tools and Identity for outbound OAuth tokens.
  • CodePipeline releases, CodeBuild builds and tests, CodeConnections links external Git, and SageMaker Pipelines runs the ML workflow; CodeDeploy shifts traffic for inference on Lambda or ECS, while SageMaker AI endpoints use their own guardrails.
  • Docker builds in CodeBuild need privileged mode; ECR pushes need service-role permissions; a full-clone source needs codeconnections:UseConnection on the CodeBuild role too.
  • Connections created by CLI or CloudFormation stay PENDING until completed in the console; V2 push triggers filter by tag or by branch and file path, pull-request triggers by branch, file path and event type.
  • Blue/green needs a full second fleet (all at once, canary, linear); rolling updates replace capacity in batches for quota-limited fleets. Auto-rollback needs alarms in AutoRollbackConfiguration plus baking periods long enough to evaluate them.
  • Spot training keeps progress only with S3 checkpoints; MaxWaitTimeInSeconds covers waiting plus running and must be at least MaxRuntimeInSeconds.
  • Batch transform: SplitType splits, BatchStrategy packs, AssembleWith formats output, and InputFilter/JoinSource/OutputFilter reshape records in the same job.
  • Gate model quality with a ConditionStep and a FailStep; test deployed endpoints after the staging deploy; use a manual approval action for human sign-off.
  • Retrain from EventBridge schedules or from CloudWatch alarm state changes on Model Monitor metrics, never from every monitoring run; start releases from Model Package State Change events.
  • Knowledge base sync is incremental and includes deletions; a new embeddings model means a new knowledge base, a full ingestion and an ID switch.
  • Call prompts, custom models and agents through stable versioned identifiers: a versioned prompt ARN, a deployment or provisioned model ARN, a pinned AgentCore endpoint.
  • CodeDeploy releases need a deployment group, an AppSpec file, a canary, linear or all-at-once deployment configuration, and alarms on the deployment group for automatic rollback.

Operating, Monitoring, and Securing ML and AI Solutions

24% of the exam
  • Bedrock invocation logging is off by default; the Model Invocations view's per-request table needs it delivered to CloudWatch Logs, and bodies over 100 KB need an S3 large-data location.
  • TimeToFirstToken measures a slow start in streaming; InvocationLatency measures the whole response. Early failures appear only in Bedrock metrics without the ModelId dimension.
  • Every Model Monitor type needs data capture and a baseline; model quality also needs ground truth joined by InferenceId and a baseline built from predictions plus labels.
  • Data quality watches inputs, model quality watches accuracy, bias drift watches outcome gaps between groups, feature attribution drift watches the SHAP importance ranking.
  • Alarm on the drift metric, then let an EventBridge rule on the alarm state start the retraining pipeline; use anomaly detection bands for seasonal metrics and Logs anomaly detection for unpredictable log patterns.
  • Production variants expose users to the new model by weight; shadow variants never do. UpdateEndpointWeightsAndCapacities changes weights in place, TargetVariant pins a single request, InvokedProductionVariant attributes outcomes.
  • Bedrock judge-based and RAG evaluations can score your own logged responses without re-invoking the model; retrieve-only jobs isolate retrieval quality.
  • Handled tool errors are invisible to Lambda Errors; emit a custom metric. A stream without messageStop or with stopReason max_tokens is incomplete, whatever the HTTP status.
  • AgentCore online evaluation samples live sessions continuously; pick session, trace or tool-level evaluators to match the failure, and alarm on the Bedrock-AgentCore/Evaluations metrics.
  • Pick the instance family that supplies the bottleneck resource: c for CPU-bound, r for memory-heavy, GPU only when the model uses it, and check GPU memory (16, 24 or 48 GB) before LLM fit.
  • Inference Recommender default jobs shortlist instances; advanced load test jobs take your traffic phases, endpoint configurations and percentile-latency stopping conditions.
  • ModelLatency plus OverheadLatency is the time inside SageMaker AI; the rest of a slow request is found with X-Ray active tracing, whose sampling is decided at the entry service.
  • AgentCore traces need the ADOT SDK and CloudWatch Transaction Search enabled once per account; metrics appear without it, spans do not.
  • Dashboards: plot percentiles for latency, use cross-account observability for one central view, and SEARCH expressions so new resources appear automatically.
  • Tags reach Cost Explorer only after activation; per-team Bedrock attribution uses tagged application inference profiles; Budgets forecast alerts warn early and Budgets actions apply IAM policies or SCPs.
  • SageMaker AI commitments are SageMaker Savings Plans (not Compute Savings Plans or EC2 RIs), sized to the steady hourly baseline.
  • Bedrock Flex discounts latency-tolerant synchronous calls; batch discounts S3 jobs; Priority costs more; idle GPU endpoints are better replaced by per-token on-demand where the model is offered.
  • AgentCore Runtime bills CPU only while processing but memory until the session ends, so stop sessions and shorten the idle timeout.
  • Bedrock reserves input plus maximum output tokens against the quota; cache checkpoints go after the static prefix and must meet the model's minimum size; resync incrementally and shrink vectors to cut RAG cost.
  • Any registry-side image scan runs after the push; a 'never stored' rule needs the Inspector SBOM Generator and Scan API inside the build, and 'keep re-checking old images' needs Inspector continuous enhanced scanning.
  • CodeGuru Reviewer is closed to new repository associations and CodeGuru Security is discontinued; Amazon Inspector code security provides SAST, SCA and IaC scanning for GitHub and GitLab.
  • A SageMaker AI job runs with its execution role; the caller needs iam:PassRole scoped to that role ARN with iam:PassedToService = sagemaker.amazonaws.com.
  • Cross-account or SSE-KMS artifacts need kms:Decrypt in both the key policy and the role's identity policy; cross-account model packages also need a model package group resource policy.
  • Prevent non-compliant jobs with explicit Deny statements on SageMaker AI condition keys, one per requirement; negated operators match a missing key, positive ones need IfExists or a Null check.
  • Network isolation blocks all outbound calls from the container, even inside a VPC, but input channels and output upload still work; a VPC without an S3 endpoint or NAT gives timeouts, not AccessDenied.
  • SageMaker InvokeEndpoint and Bedrock InvokeAgent, Retrieve and ApplyGuardrail are CloudTrail data events (off by default); CloudTrail never records prompts, which need Bedrock model invocation logging.
  • Applications on AWS should use IAM role credentials; short-term Bedrock API keys (at most 12 hours) suit key-only tools, and long-term keys tied to IAM users are for exploration only.
  • Match the guardrail policy to the need: denied topics for subjects, content filters for harm categories, input-tagged prompt attack filters for jailbreaks, sensitive information and regex filters to block or mask PII.
  • Enforce a guardrail with bedrock:GuardrailIdentifier using the versioned ARN and an explicit Deny, and keep that role away from RetrieveAndGenerate and InvokeAgent.

Keep studying MLA-C02

MLA-C02 flashcards: your questions

Yes — the whole MLA-C02 deck is free and needs no account, as is every MLA-C02 revision lesson. Create a free account to save progress and drill practice questions; full-length mock exams are a one-time unlock per certification.

Source

Exam structure, domain weights and scoring on this page come from the official MLA-C02 exam guide.