MLOps CI/CD Pipelines, Model Versioning and Rollback on AWS (MLA-C02)
MLOps CI/CD on AWS is the automation that takes ML models, prompts, fine-tuned foundation models, agents and knowledge bases from a code or data change to production safely: AWS CodePipeline, CodeBuild and CodeConnections run the release, Amazon SageMaker Pipelines runs training and evaluation, SageMaker Model Registry, MLflow and Amazon Bedrock Prompt Management keep versions, and deployment guardrails roll bad releases back. Task 3.3 of the MLA-C02 exam tests how these pieces are configured, how they are triggered, where each test gate belongs and why a release fails. This lesson follows a release from source commit to production and back again on rollback.
On this page10 sections
- What does a CI/CD pipeline for ML and AI workloads look like on AWS?
- How do you configure and troubleshoot CodePipeline, CodeBuild and CodeConnections?
- How does SageMaker Pipelines orchestrate training and inference jobs?
- How do you configure training and batch inference jobs in automated pipelines?
- How do you build automated testing into ML and AI pipelines?
- Which deployment strategies and rollback options protect a SageMaker AI endpoint update?
- How do you trigger re-training and keep knowledge bases fresh automatically?
- How do SageMaker Model Registry and MLflow version models for repeatability and audit?
- How do you manage prompts and test models and prompts before release?
- How do you automate deployment of fine-tuned foundation models and agents?
- Configure and troubleshoot CodePipeline, CodeBuild, CodeConnections, CodeCommit and CodeDeploy for ML workloads, including privileged mode, service-role permissions, pending connections, trigger filters and CodeDeploy rollback.
- Orchestrate ML workflows with SageMaker Pipelines and configure training jobs (spot, checkpoints, stopping conditions) and batch transform jobs for automation.
- Place automated tests where they belong: unit tests, metric gates with condition and fail steps, staging integration tests, test reports and manual approvals.
- Choose blue/green (all at once, canary, linear) or rolling endpoint updates and configure alarm-based and pipeline-stage rollback.
- Build re-training and knowledge base refresh triggers with EventBridge, CloudWatch alarms and ingestion jobs.
- Version models, prompts, fine-tuned foundation models and agents with Model Registry, MLflow, Prompt Management, Provisioned Throughput and AgentCore Runtime endpoints.
- Test foundation models, prompts and RAG retrieval automatically with Amazon Bedrock evaluation jobs.
What does a CI/CD pipeline for ML and AI workloads look like on AWS?
An MLOps CI/CD setup on AWS is usually two cooperating pipelines: a model-build pipeline that prepares data, trains, evaluates and registers a model (typically Amazon SageMaker Pipelines), and a release pipeline that takes an approved model, prompt or agent version and deploys it through staging to production (typically AWS CodePipeline). Each tool in the AWS developer tool chain has one job, and most troubleshooting questions start with knowing which tool owns the failing step.
| Service | Job in an ML pipeline | What it is not |
|---|---|---|
| AWS CodePipeline | Release orchestrator: source, build, test, approval and deploy stages; V2 pipelines add triggers with filters and stage rollback | Not an ML workflow engine; it starts and waits on other services |
| AWS CodeBuild | Managed, short-lived build host driven by a buildspec: unit tests, Docker builds, pushes to Amazon ECR, scripts that call SageMaker AI or Amazon Bedrock APIs, integration tests | Not a long-running server or a deployment service |
| AWS CodeConnections | Authorised link to GitHub, GitLab or Bitbucket used by source actions (and by CodeBuild for full clones) | Not a repository itself |
| AWS CodeCommit | AWS-hosted Git repository; as a pipeline source it starts the pipeline through an Amazon EventBridge rule on repository changes | Not needed when code already lives in GitHub, GitLab or Bitbucket |
| AWS CodeDeploy | Deploys to Amazon EC2/on-premises servers, AWS Lambda and Amazon ECS with traffic shifting and alarm-based rollback, for example inference code served from Lambda or ECS | Not used for SageMaker AI endpoints, which shift traffic with their own deployment guardrails |
| Amazon SageMaker Pipelines | Serverless ML workflow DAG: processing, training, tuning, evaluation, condition, registration, transform | Not a source-control-triggered release tool |
| AWS Step Functions / Amazon MWAA | General workflow orchestration (state machines / Apache Airflow DAGs) that can call ML services | No built-in ML lineage view in SageMaker Studio |
The deployment target decides the deploy action. A SageMaker AI endpoint is usually updated by an AWS CloudFormation deploy action (or a CodeBuild step calling the SageMaker AI API), and the safety of that update comes from the endpoint's own deployment guardrails, covered below. Choosing the hosting option itself (real-time, serverless, asynchronous, multi-model) belongs to Task 3.1, and building the containers and stacks belongs to Task 3.2.
How do you configure and troubleshoot CodePipeline, CodeBuild and CodeConnections?
Most pipeline failures trace back to four settings: the CodeBuild project's environment (privileged mode), its service role permissions, the CodeConnections connection status and permissions, and the pipeline's trigger configuration. Read the error and the stage that failed; together they usually point to exactly one of these.
CodeBuild projects and buildspecs
A buildspec.yml defines phases (install, pre_build, build, post_build), plus artifacts and reports sections. A typical ML build runs pytest in pre_build, runs docker build in build, and logs in to Amazon ECR and pushes the image in post_build.
- Docker builds need privileged mode. Managed CodeBuild images include Docker, but the Docker daemon only runs when the project's environment has privileged mode enabled. The symptom is "Cannot connect to the Docker daemon at unix:///var/run/docker.sock" while non-Docker commands in the same build succeed.
- ECR push needs service-role permissions. The CodeBuild service role needs
ecr:GetAuthorizationTokento log in and the layer-upload andecr:PutImageactions on the repository to push. Missing permissions show up as an authorisation error at login or push, not as a daemon error. - Architecture follows the build host. An image built on an x86_64 host without a platform flag is an x86_64 image. Build for another architecture with
docker buildx build --platformor use an ARM-based CodeBuild environment. A bigger compute type only changes speed. - Network settings (running the project in a VPC with a NAT gateway or VPC endpoints) matter only when the build must reach private resources or the internet from private subnets.
CodeConnections sources
A connection created with the AWS CLI or AWS CloudFormation starts in PENDING status. Only the console can complete the handshake with the provider: open the connection, choose Update pending connection and install or select the provider app, after which it becomes AVAILABLE. The handshake is an interactive authorisation with the provider, which is why it cannot be finished from the API.
| Source output format | Who uses the connection | Permissions needed |
|---|---|---|
| CodePipeline default (zip) | CodePipeline downloads the code and passes a zip through the S3 artifact bucket | codeconnections:UseConnection on the pipeline service role |
| Full clone | CodePipeline passes repository metadata; CodeBuild clones the repository itself, keeping Git history such as the commit ID | codeconnections:UseConnection on the pipeline role and on the CodeBuild service role |
The rule behind the table: whichever service actually opens the connection needs UseConnection on its own role, so check which service fails to reach the repository before changing permissions.
Starting the pipeline only when it should run
CodePipeline V2 pipelines with a CodeConnections source support triggers with filters: push triggers filter by Git tags, or by branches and file paths (include and exclude globs such as training/**); pull-request triggers filter by branches, file paths and pull-request event type (opened, updated, closed). In a monorepo, a branch filter plus a file-path include filter keeps front-end commits from starting an expensive training run. Filters decide whether an execution starts at all; logic inside a stage only runs after the execution, and its cost, has already begun.
CodeCommit sources
An AWS CodeCommit source action is change-detected by an Amazon EventBridge rule on the repository and branch, which the console creates for you (CloudFormation and CLI setups create it themselves and set PollForSourceChanges to false). Two roles are involved: the rule needs a role allowed to call codepipeline:StartPipelineExecution on the pipeline, and the pipeline service role needs CodeCommit read actions such as codecommit:GetBranch, codecommit:GetCommit and codecommit:UploadArchive to fetch the source. A pipeline that never starts after a push points at the rule or its role; a source stage that starts and then fails points at the pipeline role.
How does SageMaker Pipelines orchestrate training and inference jobs?
SageMaker Pipelines is the serverless ML workflow service of SageMaker AI: you define a directed acyclic graph of steps in the SageMaker Python SDK, start it with parameters, and SageMaker AI runs each step as a managed job and records lineage between data, jobs and model versions that SageMaker Studio displays. It suits teams that want ML-specific steps and lineage without running an orchestrator. Step Functions and Amazon MWAA can also orchestrate ML jobs and fit when the workflow spans many non-ML services or the team already runs Airflow, but they do not give the Studio lineage view.
Common step types: Processing (data preparation and evaluation), Training, Tuning, Model (create or register a model), Transform (batch transform), Condition (branch on a value), Fail (end the run as Failed with a message), Lambda and Callback (call out to other systems), plus QualityCheck and ClarifyCheck for baselines and bias checks.
| Setting | What it does | Use it when |
|---|---|---|
| Pipeline parameters | Typed inputs (instance type, S3 URI, threshold) supplied at start time | The same definition runs for different data, environments or schedules |
Step caching (CacheConfig with an expiry) | Reuses the outputs of an earlier successful run of a step with identical arguments instead of launching a new job | An expensive, unchanged step (for example a long processing job) would otherwise rerun while you iterate on later steps |
| Retry policy | Retries a step on listed errors such as throttling or capacity errors | Transient service errors, not repeated work |
| Parallelism configuration | Caps how many steps run at the same time | Protecting quotas or downstream systems |
Pipelines are started by people, by schedules or by events (see re-training below), and the release pipeline usually starts after a model version is registered and approved, not after a training run simply finishes.
How do you configure training and batch inference jobs in automated pipelines?
Automated training and inference jobs are configured for cost, time limits and output shape: managed spot training with checkpoints and stopping conditions for training, and batch transform settings for large offline scoring. These are settings on the job, so the pipeline step simply carries them.
Training jobs: spot, checkpoints and stopping conditions
- Managed spot training runs on spare capacity at a lower price; an interruption stops the container and SageMaker AI resumes the job when capacity returns.
- Checkpoints keep progress across interruptions: set a checkpoint S3 URI on the job and have the script save to, and load from, the local checkpoint directory (by default
/opt/ml/checkpoints). SageMaker AI syncs that directory with S3. Without this, a resumed job starts again from the first epoch. - Stopping conditions:
MaxRuntimeInSecondslimits actual training time;MaxWaitTimeInSecondslimits total elapsed time including waiting for spot capacity, and must be greater than or equal to the runtime. To meet an end-to-end deadline, set the wait to the deadline and the runtime below it with headroom above the expected training time. - Managed warm pools keep provisioned infrastructure for a keep-alive period so the next job starts faster. They shorten start-up between consecutive jobs; they do not save training progress or avoid repeated work.
Batch inference: transform steps and their settings
For scheduled scoring of a large dataset already in Amazon S3 with no idle cost, a Transform step runs a batch transform job: instances start, process the S3 input, write results to S3 and shut down. Endpoints are built for request traffic: scoring a whole dataset through one means managing an endpoint resource and sending every record as its own invocation, while batch transform starts, scores and shuts down on its own.
| Setting | Controls | Typical value |
|---|---|---|
SplitType | How an input file is split into records | Line for CSV or JSON Lines; RecordIO only for RecordIO data; unset sends the whole object |
BatchStrategy | Records per request | MultiRecord packs as many records as fit under MaxPayloadInMB; SingleRecord sends one |
MaxPayloadInMB | Per-request size cap | A capped value; a multi-gigabyte unsplit file can never fit |
AssembleWith | How responses are written to the output object | Line writes each returned record on its own line; unset concatenates them |
MaxConcurrentTransforms | Parallel requests per instance | Throughput tuning; it does not change payload size |
Data processing settings let one job reshape records without a post-processing step: InputFilter (JSONPath) removes fields such as an ID column before the model sees them, JoinSource=Input joins the original input record to its prediction, and OutputFilter selects which fields of the joined record to keep. Without the join, a field dropped by the input filter is gone from the output.
How do you build automated testing into ML and AI pipelines?
Automated testing in ML CI/CD happens at several gates, each placed where its evidence exists: code tests before the build, metric gates inside the model-build pipeline, integration tests after a staging deploy, and a human approval before production when policy requires it. A failed gate must stop the pipeline visibly, not silently.
| Gate | Where it runs | Mechanism |
|---|---|---|
| Unit tests on feature code and inference handlers | CodeBuild, before the image is built | pytest in the buildspec; a non-zero exit fails the action |
| Model metric threshold | SageMaker Pipelines, after evaluation | The evaluation step writes a JSON report declared as a PropertyFile; a ConditionStep reads the metric with JsonGet, runs registration if it passes and a FailStep with a message if not |
| Integration and latency tests | CodeBuild test action after the staging deploy | Replays recorded requests against the staging endpoint; fails on errors or slow responses |
| Kept test results | CodeBuild report groups | Tests write JUnit XML (or another supported format) and the buildspec reports section points at the files, creating a test report per run |
| Human sign-off | CodePipeline, between staging and production | A manual approval action pauses the pipeline, can notify reviewers through Amazon SNS, records who approved or rejected, and continues or fails the run |
Two principles decide most gate designs. First, a gate only works if failure is visible: a rejected model has to end the execution with a Failed status, which is the job of the FailStep, so that anything watching the pipeline sees the rejection. Second, a test belongs where its evidence exists: a claim about how the deployed endpoint behaves (errors, latency, request format) can only be checked against a deployed endpoint, so it runs after the staging deploy. Production monitoring and capacity sizing answer different questions from a pass/fail release decision, and the managed gate mechanisms above need no extra code to maintain. Scanning code and images for vulnerabilities in the pipeline belongs to Task 4.3.
Which deployment strategies and rollback options protect a SageMaker AI endpoint update?
SageMaker AI deployment guardrails update a real-time endpoint either blue/green (build a complete new fleet, then shift traffic) or rolling (replace capacity in batches), and both can roll back automatically when CloudWatch alarms listed in the endpoint's AutoRollbackConfiguration fire. You choose them in the DeploymentConfig of UpdateEndpoint or of the endpoint resource in CloudFormation.
| Mode | How traffic moves | Fits when |
|---|---|---|
| Blue/green, all at once | 100% to the new (green) fleet in one step, then a baking period | Simple updates where a short full exposure is acceptable |
| Blue/green, canary | A small canary slice first, a baking period, then all remaining traffic in one step | You want a limited blast radius before a single cut-over |
| Blue/green, linear | Several equal-sized steps, with a baking period after each | You want gradual exposure across many steps |
| Rolling | New capacity is provisioned in batches (MaximumBatchSize), with a wait interval between batches, while old capacity is removed | Large or expensive fleets where quota or cost rules out running two full fleets at once |
How auto-rollback works. SageMaker AI watches the alarms during each baking period (or wait interval); if any alarm enters ALARM, traffic returns to the old fleet. Rollback therefore needs both pieces: alarms on meaningful metrics of the new variant (for example Invocation5XXErrors or ModelLatency) and baking periods long enough for those alarms to evaluate. With no alarms configured, traffic shifting completes regardless of errors, because the alarms are the only rollback trigger. Every blue/green mode needs a full green fleet, so only a rolling update respects a tight instance quota. Shadow testing (mirroring traffic to a shadow variant) validates a candidate before promotion; it is not a rollback mechanism, and comparing shadow variants with production variants is part of Task 2.3.
Pipeline-level rollback is separate. In CodePipeline V2, a stage can be configured to roll back on failure, which reruns that stage with the source revisions of the last successful pipeline execution. Endpoint guardrails protect a single update in flight; stage rollback recovers the pipeline after a failed deploy.
AWS CodeDeploy for inference on Lambda or Amazon ECS
When inference code runs in AWS Lambda or on Amazon ECS rather than on a SageMaker AI endpoint, AWS CodeDeploy provides the traffic shifting and rollback. Its configuration has four parts:
- Application and deployment group: the deployment group names the target (a Lambda function, or an ECS service with its load balancer listeners and two target groups), the CodeDeploy service role, the deployment configuration, CloudWatch alarms and rollback settings.
- AppSpec file: for Lambda it names the function, the alias and the current and target versions; for ECS it names the task definition, container name and container port. Optional lifecycle hooks run Lambda functions as tests, such as
BeforeAllowTrafficandAfterAllowTraffic(plusAfterAllowTestTrafficon ECS, which can test the new task set through a test listener before production traffic moves). - Deployment configuration: predefined canary, linear or all-at-once shapes such as
CodeDeployDefault.LambdaCanary10Percent5Minutes,CodeDeployDefault.LambdaLinear10PercentEvery1Minute,CodeDeployDefault.ECSCanary10Percent5MinutesorCodeDeployDefault.ECSAllAtOnce, or a custom one. Lambda shifts weight between versions behind an alias; ECS blue/green shifts the load balancer listener from the original task set to the replacement task set. - Automatic rollback: enable rollback when a deployment fails and when an alarm threshold is met, and list the CloudWatch alarms (for example on the function's errors or the target group's 5XX count) on the deployment group. CodeDeploy stops the deployment and routes traffic back to the last good version.
In CodePipeline, Lambda deployments usually come from a CloudFormation or AWS SAM deploy action (SAM's DeploymentPreference creates the CodeDeploy resources), and ECS blue/green uses the Amazon ECS (Blue/Green) deploy action, which reads appspec.yaml and taskdef.json from the build artifact. Common failures: a hook function that never calls PutLifecycleEventHookExecutionStatus, so the deployment times out; a container name or port in the AppSpec that does not match the task definition; a CodeDeploy service role without permissions on the target; and alarms that are never listed on the deployment group, so nothing triggers a rollback.
How do you trigger re-training and keep knowledge bases fresh automatically?
Re-training and knowledge base refresh are both event-driven automation: Amazon EventBridge starts the right job on a schedule or when a meaningful signal occurs, and the job itself (a SageMaker AI pipeline, or a Bedrock ingestion job) does the work. The design question is always which event carries the signal you actually care about.
Re-training mechanisms
- Schedule: an EventBridge schedule (EventBridge Scheduler or a scheduled rule) can target a SageMaker AI pipeline directly and pass pipeline parameters, so a monthly retrain needs no code.
- Drift: with CloudWatch metrics enabled, SageMaker Model Monitor publishes per-feature drift metrics. A CloudWatch alarm on the metric you care about enters ALARM only when a threshold is crossed, and the alarm state change is an EventBridge event whose rule can start the pipeline. The event therefore carries the decision itself, which is the property a retraining trigger needs: an event that occurs on every monitoring run, whatever it found, carries no decision.
- New data or approval events: S3 object events (through EventBridge) can start a pipeline when new training data lands, and model package state changes can start a release (see Model Registry below).
Configuring the drift detection itself belongs to Task 4.1; this task is about wiring its signal to the retraining pipeline.
Knowledge base refresh cycles
An Amazon Bedrock knowledge base answers from embeddings in its vector store, not from the live S3 objects, so changes in the source reach the chatbot only after a sync: an ingestion job started with StartIngestionJob for a data source. Sync is incremental: it processes documents added, modified or deleted since the last sync. Automate it either on a schedule or from events, and make sure the event pattern covers every change type; a rule that matches only S3 Object Created events never syncs after deletions, so withdrawn documents stay retrievable.
Some changes cannot be made in place. The embeddings model and vector store configuration are fixed when a knowledge base is created, and vectors from different models or dimensions are not comparable. To change embeddings model with no downtime, create a new knowledge base with the new model and a vector index sized for its dimension, run ingestion jobs to populate it while the old one keeps serving, then switch the knowledge base ID in the application's configuration (switching back is the rollback). Chunking, parsing and vector store design are taught in Task 3.2.
How do SageMaker Model Registry and MLflow version models for repeatability and audit?
SageMaker Model Registry is the system of record for deployable model versions: a model package group holds every version of one model, and each model package records the artifact, inference specification, metrics, lineage to the training job and data, and an approval status (PendingManualApproval, Approved, Rejected). MLflow on SageMaker AI is the experiment-tracking layer that makes each run reproducible. Together they answer an auditor's question: which code, data and parameters produced the model in production, and who approved it.
Approval as the release trigger
When a model package's approval status changes, SageMaker AI emits a SageMaker Model Package State Change event to EventBridge. A rule matching the model package group and the Approved status can start the CodePipeline release directly, with no polling. The release trigger must carry the approval decision itself, and the state-change event is the only point in the flow where that decision exists; everything earlier in the model-build pipeline happens before anyone has reviewed the version. The same registry pattern applies to a foundation model fine-tuned with SageMaker AI training jobs and hosted on a SageMaker AI endpoint: it is still a SageMaker model, so each fine-tune is a new version in a group and only Approved versions are deployed.
| Tool | What it records | Role in the release |
|---|---|---|
| Model Registry | Versions, metrics, lineage, approval status and approval events | The catalog that deployment reads from and that governance approves in |
| MLflow on SageMaker AI | Runs: parameters, metrics, artifacts and code/data references through the open-source MLflow API; its own MLflow model registry | Reproducibility during development; feeds Model Registry when synced (below) |
A complete audit answer needs the artifact, its metrics, its lineage and its approval decision on one record, which is what a model package holds. Documentation, experiment grouping and file versioning each capture part of that story and complement the registry.
MLflow on SageMaker AI
SageMaker AI hosts MLflow for you (MLflow tracking servers and the newer MLflow Apps), so notebooks log each run with the standard MLflow client while AWS operates the infrastructure. When Model Registry sync is enabled on the tracking server or app, every model registered in MLflow automatically creates a model package group (if needed) and a model package version in SageMaker Model Registry, with no notebook changes. The setting lives on the MLflow server, and its service role needs permissions such as sagemaker:CreateModelPackageGroup and sagemaker:CreateModelPackage.
How do you manage prompts and test models and prompts before release?
Amazon Bedrock Prompt Management stores prompts as managed resources with variables, variants and immutable numbered versions, and Amazon Bedrock evaluation jobs score models, prompts and knowledge bases automatically, so a pipeline can treat a prompt like any other versioned, tested artifact. Because a managed prompt is a resource with its own ARN, the application invokes it directly, and a wording change becomes a new prompt version rather than a code release.
Drafts, versions and the prompt ARN
- A prompt has a working draft that changes every time someone saves an edit. Variables (placeholders such as
{{document}}) are filled at call time, and variants let you compare wording or models side by side in the console. - Creating a version takes an immutable snapshot of the prompt text, model and inference settings, numbered 1, 2, 3 and so on.
- An application runs a prompt by passing its ARN as the
modelIdin Converse (or InvokeModel) withpromptVariables. The ARN without a version suffix runs the draft, so experiments leak into production; the ARN ending in:<version>runs that snapshot no matter how the draft changes. - When Converse is called with a prompt ARN, the prompt's stored configuration is used and request-level overrides of that configuration are rejected. Any setting change, such as maximum tokens, is therefore a prompt change: edit the prompt, create a new version and point the application at that version's ARN.
Evaluation jobs as automated gates
A CodeBuild step can start an Amazon Bedrock evaluation job, wait for it, read the scores from S3 and fail the stage when they drop. Pick the job type by what you need to measure and whether people are available.
| Evaluation type | What it scores | Use when |
|---|---|---|
| Automatic (programmatic metrics) | Accuracy, robustness and toxicity with algorithmic metrics, on built-in or custom datasets | Task metrics that a formula can compute |
| LLM-as-a-judge model evaluation | An evaluator model rates responses for qualities such as correctness, completeness, helpfulness and harmfulness, on your prompt dataset | Subjective quality must be scored with no human reviewers |
| Human evaluation | A work team rates responses | Reviewers are available and their judgement is required |
| RAG evaluation, retrieve only | The knowledge base's retrieved passages: context relevance and (with ground-truth answers) context coverage | Isolating retrieval quality, for example after a chunking change |
| RAG evaluation, retrieve and generate | The generated answer as well: correctness, completeness, helpfulness, faithfulness and similar | Judging the end-to-end RAG answer |
Model evaluation jobs call the generator model directly and bypass the knowledge base, so they cannot measure retrieval. Runtime monitoring such as CloudWatch generative AI observability watches production behaviour and belongs to Task 4.1; it is not a pre-release test.
How do you automate deployment of fine-tuned foundation models and agents?
Fine-tuned foundation models and agents are released the same way as models: each build produces an immutable version, a stable identifier the application calls is pointed at a tested version, and promotion or rollback moves that pointer instead of changing application code.
Fine-tuned models in Amazon Bedrock
A Bedrock model customization job produces a custom model, but the bare custom model ARN is not something an application invokes. Serve it one of two ways:
| Option | Billing | Model ID the application calls |
|---|---|---|
| Custom model deployment (on-demand inference, for supported models such as Amazon Nova) | Pay per use, no reserved capacity | The deployment ARN |
| Provisioned Throughput | Reserved model units, with or without a commitment term | The provisioned model ARN |
To release a newer fine-tune without buying capacity or changing the application, update the existing Provisioned Throughput with UpdateProvisionedModelThroughput: its desiredModelId can be the base model or another custom model customized from the same base model, and the provisioned model ARN stays the same. Amazon Bedrock Custom Model Import is for models trained outside Bedrock (Task 3.1), and a fine-tune hosted on SageMaker AI is versioned in Model Registry, as above.
Agents on Amazon Bedrock AgentCore Runtime
- Every update to an agent runtime (for example a new container image URI) creates a new immutable version.
- The DEFAULT endpoint automatically points at the latest version, so an application calling it receives every update as soon as the pipeline deploys it.
- A custom endpoint (for example
production) is pinned to a chosen version and moves only whenUpdateAgentRuntimeEndpointis called. A pipeline deploys, tests the new version, then promotes by pointing the production endpoint at it; rollback points it back at the previous version. - AgentCore Runtime requires linux/arm64 container images. An image built on an x86_64 CodeBuild host without a platform flag pushes fine and passes local tests, but the new runtime version fails to start.
The original Amazon Bedrock Agents (now Agents Classic, with agent versions and aliases) is closed to new customers, so new agent release pipelines are built on AgentCore. Choosing agent hosting and protocols belongs to Task 3.1, and agent state and infrastructure to Task 3.2.
Tip. Task 3.3 tests whether you can configure, trigger and troubleshoot automated ML and generative AI release workflows on AWS: setting up and fixing CodePipeline, CodeBuild, CodeConnections, CodeCommit and CodeDeploy; orchestrating training, evaluation and batch inference with SageMaker Pipelines; placing unit tests, metric gates, integration tests and approvals in the right stage; choosing endpoint deployment strategies and rollback; triggering retraining and knowledge base refreshes from events; and versioning models, prompts, fine-tuned foundation models and agents so releases are repeatable and reversible. Questions are scenarios with a stated requirement or a failure symptom, and some are multiple response.
- CodePipeline releases, CodeBuild builds and tests, CodeConnections links external Git, and SageMaker Pipelines runs the ML workflow; CodeDeploy shifts traffic for inference on Lambda or ECS, while SageMaker AI endpoints use their own guardrails.
- Docker builds in CodeBuild need privileged mode; ECR pushes need service-role permissions; a full-clone source needs codeconnections:UseConnection on the CodeBuild role too.
- Connections created by CLI or CloudFormation stay PENDING until completed in the console; V2 push triggers filter by tag or by branch and file path, pull-request triggers by branch, file path and event type.
- Blue/green needs a full second fleet (all at once, canary, linear); rolling updates replace capacity in batches for quota-limited fleets. Auto-rollback needs alarms in AutoRollbackConfiguration plus baking periods long enough to evaluate them.
- Spot training keeps progress only with S3 checkpoints; MaxWaitTimeInSeconds covers waiting plus running and must be at least MaxRuntimeInSeconds.
- Batch transform: SplitType splits, BatchStrategy packs, AssembleWith formats output, and InputFilter/JoinSource/OutputFilter reshape records in the same job.
- Gate model quality with a ConditionStep and a FailStep; test deployed endpoints after the staging deploy; use a manual approval action for human sign-off.
- Retrain from EventBridge schedules or from CloudWatch alarm state changes on Model Monitor metrics, never from every monitoring run; start releases from Model Package State Change events.
- Knowledge base sync is incremental and includes deletions; a new embeddings model means a new knowledge base, a full ingestion and an ID switch.
- Call prompts, custom models and agents through stable versioned identifiers: a versioned prompt ARN, a deployment or provisioned model ARN, a pinned AgentCore endpoint.
- CodeDeploy releases need a deployment group, an AppSpec file, a canary, linear or all-at-once deployment configuration, and alarms on the deployment group for automatic rollback.
Frequently asked questions
Where does AWS CodeDeploy fit in an ML release pipeline?
CodeDeploy releases inference code that runs on AWS Lambda, Amazon ECS or Amazon EC2: a deployment group names the target, an AppSpec file names the versions or task definition, a deployment configuration sets canary, linear or all-at-once traffic shifting, and CloudWatch alarms on the deployment group trigger automatic rollback. SageMaker AI endpoints are not CodeDeploy targets; they use their own deployment guardrails with alarms in the endpoint's AutoRollbackConfiguration.
What is the difference between canary and linear traffic shifting on a SageMaker AI endpoint?
Both are blue/green modes that build a full new fleet first. Canary shifts a small slice of traffic, waits a baking period while alarms are watched, then shifts all the rest in one step. Linear shifts traffic in several equal steps with a baking period after each. Neither rolls back unless CloudWatch alarms are configured for auto-rollback.
What does a CodeBuild project need to build and push Docker images?
Privileged mode enabled in the project's environment, because the Docker daemon only starts in a privileged build environment (without it the build reports that it cannot connect to the Docker daemon), and a service role with ecr:GetAuthorizationToken plus the layer-upload and ecr:PutImage actions on the target Amazon ECR repository.
How do I trigger a SageMaker AI retraining pipeline when data drift is detected?
Enable CloudWatch metrics on the SageMaker Model Monitor schedule, create a CloudWatch alarm on the drift metric that matters, and create an Amazon EventBridge rule on that alarm's state change to ALARM that targets the SageMaker AI pipeline. The alarm changes state only when the threshold is crossed, so retraining runs only when drift actually matters.
How do I stop prompt edits in Amazon Bedrock Prompt Management from changing production?
Create a version of the prompt and have the application call the prompt ARN that ends with that version number. The ARN without a version suffix runs the working draft, which changes every time someone saves an edit; a version is an immutable snapshot of the prompt text, model and inference settings.
How do I version and roll back an agent on Amazon Bedrock AgentCore Runtime?
Every runtime update creates an immutable version, and the DEFAULT endpoint always follows the newest one. Create a custom endpoint such as production pinned to a tested version and point the application at it; promote or roll back by updating that endpoint to reference a different existing version.
How do I refresh an Amazon Bedrock knowledge base after documents change in Amazon S3?
Start an ingestion job (StartIngestionJob) for the data source after the change, on a schedule or from S3 events delivered through EventBridge. Sync is incremental and processes added, modified and deleted documents, so the trigger must also fire on deletions. Changing the embeddings model requires a new knowledge base, because the model is fixed when the knowledge base is created.
What does Model Registry sync do for MLflow on SageMaker AI?
When Model Registry sync is enabled on a managed MLflow tracking server or MLflow App, each model registered in MLflow automatically creates a model package version (and the group if needed) in SageMaker Model Registry. The MLflow server's service role needs model package permissions; notebooks do not change.
Source
This lesson covers the "Deployment and Orchestration of ML and AI Workflows" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.
- AWS Certified Machine Learning Engineer – Associate (MLA-C02) exam guide — Amazon Web Services
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.