SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
ACE · Domain 3

Ensuring the successful operation of a cloud solution practice questions

Ensuring the successful operation of a cloud solution is worth 30% of the ACE exam — the heaviest of the 4 domains. Managing compute, storage/data, and networking resources day-to-day, plus monitoring and logging. Official (approximate) weighting ~30%. 6 fully worked examples are further down this page, answers included.

Exam weight
30%
the heaviest of the 4 domains
Questions
80
across 4 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 4 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Ensuring the successful operation of a cloud solution questions, fully explained

Questions from the ACE bank mapped to domain 3, with the answer key and the reasoning behind every option. None of them repeat the examples on the main ACE practice page.

Question 1Ensuring the successful operation of a cloud solution

A Compute Engine VM hosts a file server on a 2 TB persistent disk. Management requires an automatic daily backup of the disk retained for 14 days, with no manual intervention. What is the most appropriate solution?

Choose one.

  • a
    Create a custom image of the disk every day using a cron job on the VM.

    Images are meant as immutable templates for creating new boot disks, not for incremental daily backups, and a self-managed cron job is exactly the manual machinery a snapshot schedule replaces.

  • b
    Enable object versioning on a Cloud Storage bucket and rsync the disk contents nightly.

    File-level rsync to Cloud Storage requires custom scripting, misses open files and disk metadata, and is not a block-level restore path like a snapshot.

  • c
    Create a snapshot schedule resource policy with a daily frequency and a 14-day retention, and attach it to the persistent disk. Correct

    Snapshot schedules are the native, fully managed way to take periodic disk snapshots and automatically expire them after a defined retention period.

  • d
    Create a machine image of the VM manually every morning.

    Machine images capture the whole VM and are useful for cloning, but taking them manually every day fails the no-manual-intervention requirement and retention would also be manual.

The concept

Snapshot schedules are Compute Engine resource policies that automatically create incremental snapshots of a persistent disk on a defined cadence and delete them after a retention window.

Why that’s the answer

A snapshot schedule attached to the disk takes the daily snapshot and enforces the 14-day retention without any scripts, cron jobs, or human action. Snapshots are incremental, so a 2 TB disk backs up efficiently after the first snapshot. Custom images and machine images are templates for provisioning rather than a rotation-managed backup mechanism, and rsync to Cloud Storage is an unmanaged, file-level workaround.

How to reason it out
  1. Create the policy: gcloud compute resource-policies create snapshot-schedule daily-backup --max-retention-days=14 --daily-schedule --start-time=03:00 --region=REGION.
  2. Attach it to the disk: gcloud compute disks add-resource-policies DISK_NAME --resource-policies=daily-backup --zone=ZONE.
  3. Verify snapshots appear daily with gcloud compute snapshots list and that ones older than 14 days are auto-deleted.

Exam tip: For automated periodic disk backups with retention, attach a snapshot schedule resource policy to the disk.

Managing Compute Engine, GKE, and Cloud Run: Day-2 Operations — the lesson that teaches this.

Question 2Ensuring the successful operation of a cloud solution

You have configured a Compute Engine VM with a hardened OS, monitoring agents, and your application runtime. The operations team must be able to launch dozens of identical new VMs from this configuration, including in other regions, as the standard build. What should you create?

Choose one.

  • a
    A snapshot of the boot disk that each engineer restores into a new disk when they need a VM.

    Snapshots are point-in-time backups; while a disk can be created from one, they are not designed as a golden-build distribution mechanism and cannot be set as the image in an instance template.

  • b
    An instance template that references the running VM by name.

    Instance templates do not reference a live VM; they need a source image or disk definition, so you must first capture the configuration as an image.

  • c
    A custom image created from the VM's boot disk, used as the source image for new instances and instance templates. Correct

    Custom images are the standard, multi-regional building block for launching identical VMs; they can be used directly in instance templates and shared across projects.

  • d
    A startup script that reinstalls the agents and runtime on each fresh public OS image.

    Re-running installation on every boot is slow, can drift as package versions change, and does not guarantee the byte-identical hardened build that an image does.

The concept

Custom images capture a configured boot disk as a reusable, geo-replicated template (a golden image) for provisioning identical Compute Engine instances anywhere.

Why that’s the answer

Creating a custom image from the prepared boot disk gives the operations team a single artifact that every new VM, MIG, or instance template can reference, in any region, with guaranteed consistency. Snapshots are backups rather than a provisioning standard, instance templates require an image (they cannot point at a running VM), and startup-script rebuilds are slow and prone to drift.

How to reason it out
  1. Stop the VM (or use --force-create carefully) so the disk is consistent.
  2. Create the image: gcloud compute images create golden-v1 --source-disk=DISK --source-disk-zone=ZONE.
  3. Reference the image in instance templates or gcloud compute instances create --image=golden-v1 for new VMs in any region.
  4. Optionally organize versions with an image family so templates always pick up the latest build.

Exam tip: Use custom images (optionally in an image family) as the golden build for launching identical VMs; snapshots are for backup.

Managing Compute Engine, GKE, and Cloud Run: Day-2 Operations — the lesson that teaches this.

Question 3Ensuring the successful operation of a cloud solution

Your team's snapshot storage bill is growing because old, manually created snapshots of decommissioned disks are never removed. You need to list all snapshots older than 90 days in the project and delete the ones no longer needed using the gcloud CLI. Which command sequence is correct?

Choose one.

  • a
    gcloud compute snapshots list --filter="creationTimestamp<'2026-05-01'" to identify them, then gcloud compute snapshots delete SNAPSHOT_NAME for each. Correct

    snapshots list with a creationTimestamp filter surfaces old snapshots, and snapshots delete removes them individually; this is the standard cleanup workflow.

  • b
    gcloud compute disks list --filter=snapshots, then gcloud compute disks delete for each stale entry.

    Snapshots are independent resources; deleting disks does not delete their snapshots, and disks list has no snapshot inventory view.

  • c
    gcloud compute images list, then gcloud compute images delete for each snapshot shown.

    Images and snapshots are separate resource types; the images commands never list or delete snapshots.

  • d
    gcloud compute snapshots list, then detach each snapshot from its disk with disks remove-resource-policies before it can be deleted.

    remove-resource-policies detaches a snapshot schedule from a disk; existing snapshots are standalone objects and are deleted directly, with no detach step.

The concept

Snapshots are standalone project resources managed with the gcloud compute snapshots command group; they survive disk deletion and accrue storage costs until explicitly deleted.

Why that’s the answer

gcloud compute snapshots list supports --filter expressions on fields such as creationTimestamp, letting you find snapshots older than a cutoff date, and gcloud compute snapshots delete removes each one. Because snapshots are decoupled from their source disks, disk- or image-level commands cannot manage them, and no detachment is required before deletion.

How to reason it out
  1. List old snapshots: gcloud compute snapshots list --filter="creationTimestamp<'2026-05-01'" --format="value(name)".
  2. Review the list to confirm none are needed for restores or compliance.
  3. Delete them: gcloud compute snapshots delete SNAP1 SNAP2 ... (accepts multiple names).
  4. Prevent recurrence by attaching snapshot schedules with --max-retention-days to active disks.

Exam tip: Old snapshots keep billing after their disks are gone — audit with snapshots list --filter and clean up with snapshots delete.

Managing Compute Engine, GKE, and Cloud Run: Day-2 Operations — the lesson that teaches this.

Question 4Ensuring the successful operation of a cloud solution

You have just been granted access to an existing GKE cluster. Before making changes, you want a quick inventory of every Pod running in the cluster, across all namespaces, including which node each Pod is scheduled on. Which command should you run after fetching cluster credentials?

Choose one.

  • a
    kubectl get nodes -o wide

    get nodes lists the cluster's nodes and their versions/IPs, but shows nothing about which Pods are running.

  • b
    kubectl get pods --all-namespaces -o wide Correct

    get pods --all-namespaces lists Pods in every namespace, and -o wide adds the node name, Pod IP, and other placement details.

  • c
    gcloud container clusters describe CLUSTER_NAME

    clusters describe returns cluster-level configuration (node pools, networking, versions) from the GKE control plane, not the live Pod workload inventory.

  • d
    kubectl describe deployment --all-namespaces

    Describing Deployments shows replica counts and events for Deployment objects only; it misses Pods from StatefulSets, DaemonSets, and Jobs and doesn't produce a per-Pod node listing.

The concept

Cluster inventory in GKE is a kubectl task: gcloud manages the cluster resource itself, while kubectl queries the Kubernetes API for live objects like nodes, Pods, and Services.

Why that’s the answer

kubectl get pods --all-namespaces (or -A) is the one command that enumerates every Pod regardless of namespace or owning controller, and the -o wide output format appends the NODE column you need. Node listings, gcloud describe output, and Deployment summaries each answer a different question and would miss Pods or their placement.

How to reason it out
  1. Fetch credentials: gcloud container clusters get-credentials CLUSTER_NAME --region=REGION.
  2. Run kubectl get pods --all-namespaces -o wide to list every Pod with its node, IP, and status.
  3. Follow up on any anomaly with kubectl describe pod POD -n NAMESPACE for events and details.

Exam tip: kubectl get pods -A -o wide is the fastest full-cluster Pod inventory, including node placement.

Managing Compute Engine, GKE, and Cloud Run: Day-2 Operations — the lesson that teaches this.

Question 5Ensuring the successful operation of a cloud solution

A new GKE Standard cluster in project team-apps cannot start Pods that use container images from an Artifact Registry repository in project shared-registry. Pods stay in ImagePullBackOff with a 403 error in the events. The image path and tag are correct. What is the most likely fix?

Choose one.

  • a
    Grant the node pool's service account the roles/artifactregistry.reader role on the shared-registry repository or project. Correct

    GKE nodes pull images using the node service account's credentials; a 403 from a repository in another project means that account lacks Artifact Registry read permission there.

  • b
    Recreate the cluster with larger nodes so the image has room to be pulled.

    Insufficient disk or memory produces eviction or disk-pressure symptoms, not a 403 permission error during the pull.

  • c
    Add an imagePullSecrets entry with a downloaded service account key to every Pod spec.

    Long-lived exported keys are discouraged and unnecessary on GKE: granting the node service account (or a Workload Identity principal) the reader role solves cross-project pulls without key files.

  • d
    Enable the Artifact Registry API in the team-apps project.

    The API needs to be enabled in the project hosting the repository, and a disabled API would surface as an API-not-enabled error, not an IAM 403 on pull.

The concept

GKE nodes authenticate to Artifact Registry as the node pool's service account. Same-project pulls usually work via the default compute service account's project roles, but cross-project repositories require an explicit Artifact Registry read grant.

Why that’s the answer

A 403 on image pull is an authorization failure: the node service account in team-apps has no roles/artifactregistry.reader binding in shared-registry. Granting that role on the repository (or the shared-registry project) lets kubelet pull the image and the Pods start. Node sizing is unrelated to a 403, imagePullSecrets with exported keys is an anti-pattern when IAM can authorize the pull, and API enablement matters in the repository's project and produces a different error.

How to reason it out
  1. Identify the node pool's service account: gcloud container node-pools describe POOL --cluster=CLUSTER --format="value(config.serviceAccount)".
  2. Grant it read access: gcloud artifacts repositories add-iam-policy-binding REPO --project=shared-registry --location=LOCATION --member=serviceAccount:NODE_SA --role=roles/artifactregistry.reader.
  3. Delete the failing Pods (or wait for backoff retry) so kubelet re-attempts the pull.
  4. Confirm with kubectl get pods that the Pods reach Running.

Exam tip: ImagePullBackOff with a 403 on GKE means the node service account lacks roles/artifactregistry.reader on the image's repository.

Managing Compute Engine, GKE, and Cloud Run: Day-2 Operations — the lesson that teaches this.

Question 6Ensuring the successful operation of a cloud solution

Your GKE Standard cluster runs general-purpose workloads on its default node pool. A data science team now needs to run ML inference Pods that require NVIDIA GPUs, without changing the existing nodes. What should you do?

Choose one.

  • a
    Edit the existing default node pool in place to attach GPUs to its current nodes.

    Machine type and accelerator configuration of a node pool cannot be changed in place; GPU capacity requires creating a node pool with the accelerator specified.

  • b
    Create a new node pool with an accelerator configuration (for example --accelerator type=nvidia-l4,count=1) and schedule the ML Pods onto it with a node selector and GPU resource requests. Correct

    GPUs are attached per node pool at creation; adding a dedicated GPU node pool leaves existing nodes untouched and lets the ML Pods target GPU nodes explicitly.

  • c
    Set nvidia.com/gpu resource limits on the Pods; GKE will automatically add GPUs to whatever node runs them.

    In a Standard cluster a GPU limit only makes Pods unschedulable until GPU nodes exist; GKE Standard does not attach hardware to nodes on demand.

  • d
    Migrate the inference Pods to Cloud Run, since GKE cannot run GPU workloads.

    GKE fully supports GPU node pools; a platform migration is unnecessary and doesn't address the stated requirement to run these Pods in the cluster.

The concept

In GKE Standard, hardware characteristics (machine type, GPUs, local SSD) are properties of a node pool. Heterogeneous workloads are served by running multiple node pools in one cluster.

Why that’s the answer

Creating a GPU node pool with --accelerator adds GPU-equipped nodes alongside the existing pool. The ML Pods then request GPUs (resources.limits nvidia.com/gpu: 1) and use a node selector such as cloud.google.com/gke-accelerator to land on those nodes. Existing workloads and nodes are unaffected. Node pools cannot have accelerators bolted on after creation, Standard clusters don't provision hardware from Pod specs alone, and moving to Cloud Run is out of scope.

How to reason it out
  1. Create the pool: gcloud container node-pools create gpu-pool --cluster=CLUSTER --accelerator type=nvidia-l4,count=1 --machine-type=g2-standard-4 --num-nodes=1.
  2. Ensure the NVIDIA device drivers are installed (GKE can auto-install them via the gpu-driver-version option).
  3. Add resources.limits with nvidia.com/gpu: 1 and a gke-accelerator node selector to the inference Pod spec.
  4. Apply the manifests and verify placement with kubectl get pods -o wide.

Exam tip: Add capabilities like GPUs to a GKE cluster by creating a dedicated node pool — node pool hardware can't be edited in place.

Managing Compute Engine, GKE, and Cloud Run: Day-2 Operations — the lesson that teaches this.

What ACE domain 3 tests, topic by topic

The official exam guide breaks Ensuring the successful operation of a cloud solution into 4 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published ACE practice questions per topic in Ensuring the successful operation of a cloud solution
TopicWhat it coversQuestions
Managing compute resourcesOfficial ACE exam-guide sub-section. Remotely connecting to Compute Engine; viewing running instances; snapshots and images (create/view/delete/schedule); viewing GKE inventory (nodes, Pods, Services); configuring GKE access to Artifact Registry; working with node pools (add/edit/remove, autoscaling); working with Kubernetes resources (Pods, Services, StatefulSets); horizontal/vertical Pod autoscaling; GKE Autopilot Pod resource requests; deploying new Cloud Run versions; traffic splitting (Cloud Run, functions, GKE); Cloud Run autoscaling; attaching GPUs/TPUs; deploying an agent to Agent Runtime; managing notebooks (Workbench, BigQuery); managing developer environments (Cloud Workstations).20
Managing storage and data solutionsOfficial ACE exam-guide sub-section. Managing and securing objects in Cloud Storage buckets; object lifecycle management policies; executing queries against data instances (Cloud SQL, BigQuery, Bigtable, Spanner, Firestore, AlloyDB); estimating storage costs; backing up and restoring database instances; reviewing job status (Dataflow, BigQuery); using Database Center to manage the database fleet; configuring customer-managed encryption keys (CMEK).20
Managing networking resourcesOfficial ACE exam-guide sub-section. Resizing a subnet's IPv4 range; reserving static external or internal IP addresses; adding custom static routes in a VPC; using Cloud DNS and Cloud NAT; managing VPC firewall rules and Cloud NGFW policies.20
Monitoring and loggingOfficial ACE exam-guide sub-section. Creating Cloud Monitoring alerts on resource metrics; creating and ingesting custom metrics; configuring audit logs (VPC Flow Logs, audit logs, firewall logs); exporting logs to external systems; configuring log buckets, log analytics, and log routers; viewing and filtering logs in Cloud Logging; using diagnostic tools (Cloud Trace, Cloud Profiler, Query Insights); Personalized Service Health; configuring Ops Agent; deploying Managed Service for Prometheus; using Gemini Cloud Assist for Monitoring; Active Assist; Cloud Hub.20
Total80

Revise Ensuring the successful operation of a cloud solution before you drill it

Other ACE domains

Ensuring the successful operation of a cloud solution: your questions

Ensuring the successful operation of a cloud solution is domain 3 of the ACE exam guide and carries 30% of the scored content — the heaviest of the 4 domains. On a 55-question paper that works out to roughly 17 questions, though Google Cloud does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official ACE exam guide.