Skip to content

feat(gke): Add G4 Confidential VM support with Blackwell GPUs and CMEK storage in cluster toolkit - #5817

Merged
shubpal07 merged 5 commits into
GoogleCloudPlatform:developfrom
shubpal07:shubham/cg-4
Jun 24, 2026
Merged

shubpal07 merged 5 commits into
GoogleCloudPlatform:developfrom
shubpal07:shubham/cg-4

Conversation

@shubpal07

Copy link
Copy Markdown
Contributor

Description

This PR introduces native support for GKE G4 Confidential VMs (CVM) equipped with NVIDIA Blackwell GPUs (RTX 6000 Ada) running in Confidential GPU mode (PCIe Secure Passthrough encryption), along with secure regional Customer-Managed Encryption Keys (CMEK) storage.

These enhancements are implemented as reusable, opt-in parameters inside the core Cluster Toolkit GKE scheduler, compute, and storage modules. This ensures 100% backward compatibility for all existing GKE blueprints while providing highly secure AI/ML infrastructure with minimal configuration.


Key Features & Enhancements

1. Core Module Enhancements (Opt-in & Reusable)

  • gke-node-pool: Added support for the dynamic confidential_nodes block in GKE node pools, enabling guest-level memory encryption (AMD SEV/SEV-SNP).
  • gke-cluster: Enabled cluster-level confidential_nodes configuration. Added a conditional system node pool machine_type toggle (defaulting to n2d-standard-16 when CVM is active) to prevent GKE control plane creation errors on CVM-enabled clusters.
  • gke-storage: Added Cloud KMS CMEK key integration and conditional StorageClass parameter rendering. Set wait_for_rollout = false for GKE storage manifests to prevent deployment hangs on deferred-binding volumes.

2. Core Framework Enhancements

  • kubectl-apply (Helm Core): Fixed a core framework bug where the Helm provider had a hardcoded atomic = true property. By linking atomic to the module's wait_for_rollout variable, we successfully allow Helm-deployed resources (like deferred-binding StorageClasses and PVCs) to deploy instantly without hanging the provisioning pipeline.

3. User-Facing Blueprint & Guides

  • examples/gke-g4-confidential: Added a production-ready, clean blueprint (gke-g4-confidential.yaml) and a customizable overrides file (gke-g4-confidential-deployment.yaml).
  • README.md: Comprehensive customer deployment guide detailing GCS state setup, regional Cloud KMS commands, a deep-dive on the Kubernetes storage billing lifecycle, and GKE cluster version requirements.

Verification & Testing

The blueprint and core module modifications have been subjected to rigorous end-to-end functional verification on Google Cloud:

E2E Validation Workloads (Passed)

Two custom, non-branded verification manifests were deployed and executed on the live confidential G4 cluster:

  1. Compute Workload (g4-verification-test.yaml):
    • Asserted guest-level AMD SEV CPU memory encryption: Memory Encryption Features active: AMD SEV (Passed)
    • Asserted G4 Blackwell GPU Confidential Computing state: CC status: ON and Confidential Compute GPUs Ready state: ready (Passed)
    • Executed PyTorch CUDA matrix multiplication workload over secure PCIe passthrough (Passed)
  2. Storage Workload (g4-verification-storage-test.yaml):
    • Dynamically provisioned and mounted a 100Gi GCE Hyperdisk Balanced volume under the custom StorageClass.
    • Verified secure dynamic disk attachment, file write/read loops, and subsequent GPU tensor math execution (Passed).
  3. Default Path Verification: Re-deployed and verified the cluster in its pristine default state (enable_confidential_storage: false), proving that the storage manifests are 100% robust and compatible across both standard-encrypted and CMEK-encrypted storage paths.

References & Cloud Documentation

Submission Checklist

NOTE: Community submissions can take up to 2 weeks to be reviewed.

Please take the following actions before submitting this pull request.

  • Fork your PR branch from the Toolkit "develop" branch (not main)
  • Test all changes with pre-commit in a local branch #
  • Confirm that "make tests" passes all tests
  • Add or modify unit tests to cover code changes
  • Ensure that unit test coverage remains above 80%
  • Update all applicable documentation
  • Follow Cluster Toolkit Contribution guidelines #

@shubpal07
shubpal07 requested a review from a team as a code owner June 19, 2026 13:53
@shubpal07 shubpal07 added the release-key-new-features Added to release notes under the "Key New Features" heading. label Jun 19, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the Cluster Toolkit by integrating support for GKE G4 Confidential VMs with Blackwell GPUs and Customer-Managed Encryption Keys for storage. These additions provide a highly secure and performant infrastructure for AI/ML workloads, ensuring data privacy and integrity through hardware-level encryption. The changes are implemented as opt-in parameters, maintaining backward compatibility while offering advanced security features.

Highlights

  • GKE G4 Confidential VM Support: Introduced native support for GKE G4 Confidential VMs (CVM) with NVIDIA Blackwell GPUs (RTX 6000 Ada) running in Confidential GPU mode, utilizing PCIe Secure Passthrough encryption.
  • Customer-Managed Encryption Keys (CMEK) Storage: Added secure regional Customer-Managed Encryption Keys (CMEK) storage integration for GKE Persistent Volumes.
  • Core Module Enhancements: Enhanced gke-node-pool to support dynamic confidential_nodes, gke-cluster for cluster-level confidential nodes and conditional system node pool machine types, and gke-storage for CMEK and wait_for_rollout = false.
  • Helm atomic Property Fix: Fixed a core framework bug in kubectl-apply where the Helm provider had a hardcoded atomic = true, now dynamically linked to wait_for_rollout for immediate deployment of deferred-binding resources.
  • New Example Blueprint and Documentation: Provided a new production-ready blueprint (examples/gke-g4-confidential) and comprehensive README.md documentation detailing setup, configuration, verification, and cleanup for the G4 Confidential cluster.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini or @gemini-code-assist . Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new blueprint and example configuration for deploying a GKE G4 Confidential Computing cluster with AMD SEV-SNP memory encryption and NVIDIA Blackwell GPUs, along with supporting updates to GKE node pool, storage, and cluster modules. Feedback on the changes identifies three critical issues: an undefined vars.labels reference in the blueprint that will cause a compilation error, a potential cluster creation failure when confidential nodes are enabled but the system node pool machine type defaults to an incompatible e2-standard-4 shape, and a non-existent CUDA image version (13.0.0) used in the nvidia-smi job template.

Comment thread examples/gke-g4-confidential/gke-g4-confidential.yaml Outdated
Comment thread modules/scheduler/gke-cluster/main.tf
Comment thread examples/gke-g4-confidential/gke-g4-confidential.yaml Outdated
@shubpal07
shubpal07 force-pushed the shubham/cg-4 branch 2 times, most recently from e14b824 to ee236ff Compare June 19, 2026 14:26
@shubpal07

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new blueprint to deploy a GKE G4 Confidential Computing cluster optimized for ML workloads using AMD SEV-SNP and NVIDIA Blackwell GPUs. It updates several core modules (gke-node-pool, gke-storage, gke-cluster, and kubectl-apply) to support confidential nodes, confidential instance types, and confidential storage with CMEK. The review feedback suggests two key improvements: adding a fail-fast validation in the storage module to ensure a KMS key is provided when confidential storage is enabled, and refining the cluster module's precondition to only check the system node pool machine type when the system node pool is actually enabled.

Comment thread modules/file-system/gke-storage/main.tf
Comment thread modules/scheduler/gke-cluster/main.tf
Comment thread examples/gke-g4-confidential/g4-verification-storage-test.yaml Outdated

@SwarnaBharathiMantena SwarnaBharathiMantena left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please check if we need to add enable-confidential-storage variable to gke-cluster and gke-node-pool modules.

Comment thread modules/compute/gke-node-pool/variables.tf
Comment thread modules/scheduler/gke-cluster/variables.tf
Comment thread examples/gke-g4-confidential/gke-g4-confidential.yaml Outdated

@kadupoornima kadupoornima left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this PR. Just a couple of points:

  • In the g4-verification-*.yaml manifests, you can change the || echo "..." fallbacks to || { echo "..."; exit 1; }. Right now, if nvidia-smi conf-compute fails, the Kubernetes Job will still report a Success status, masking the failure.
  • Regarding Swarna's comment.. while GKE does allow enabling confidential nodes only at the node pool level (without setting it on the cluster), keeping it on both is the right move for a strict Confidential blueprint. Consider exporting the values via outputs.tf as suggested to reduce boilerplate.

@shubpal07

Copy link
Copy Markdown
Contributor Author

Please check if we need to add enable-confidential-storage variable to gke-cluster and gke-node-pool modules.

Acknowledged. Thanks for bringing this up. I have exposed enable_confidential_storage as a variable in both the gke-cluster and gke-node-pool modules and mapped it directly to their resource node_config blocks. This ensures that node boot disks (for both the system pool and workload pools) are encrypted using the CVM's ephemeral keys.

@shubpal07

Copy link
Copy Markdown
Contributor Author

Thanks for this PR. Just a couple of points:

  • In the g4-verification-*.yaml manifests, you can change the || echo "..." fallbacks to || { echo "..."; exit 1; }. Right now, if nvidia-smi conf-compute fails, the Kubernetes Job will still report a Success status, masking the failure.
  • Regarding Swarna's comment.. while GKE does allow enabling confidential nodes only at the node pool level (without setting it on the cluster), keeping it on both is the right move for a strict Confidential blueprint. Consider exporting the values via outputs.tf as suggested to reduce boilerplate.

Thanks for catching it @kadupoornima!. The previous || echo fallback allowed the shell to exit with status 0, masking hardware or security faults and causing GKE to falsely report the validation Job as successful.

I have updated the shell wrappers in both manifests to exit with 1 on failure, ensuring that Kubernetes Job statuses accurately reflect any validation or hardware errors. Also I have removed the redundant, hardcoded CVM configurations from the node pool settings in the blueprint. The node pool will now dynamically and automatically inherit the cluster's confidentiality state under the hood.

@shubpal07

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new blueprint, gke-g4-confidential, which provisions a GKE cluster running on Confidential VMs with NVIDIA Blackwell GPUs and optional Confidential Storage. It updates the GKE cluster, node pool, and storage modules to support confidential nodes, confidential storage, and boot disk KMS keys. Review feedback highlights two key issues: the g4-pool node pool in the new blueprint is missing the boot_disk_kms_key setting, which will trigger a deployment precondition failure when confidential storage is enabled, and the system node pool in the GKE cluster module should explicitly configure the confidential_nodes block for consistency and to prevent API validation errors.

Comment thread examples/gke-g4-confidential/gke-g4-confidential.yaml Outdated
Comment thread modules/scheduler/gke-cluster/main.tf

@SwarnaBharathiMantena SwarnaBharathiMantena left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@SwarnaBharathiMantena SwarnaBharathiMantena removed their assignment Jun 24, 2026
@shubpal07

shubpal07 commented Jun 24, 2026 •

Copy link
Copy Markdown
Contributor Author

Failing tests:

PR-test-gke-a3-highgpu -> Test triggered by babysit (Reservation not available)
PR-test-gke-a4x -> Failing on develop as well (reservation not available)
PR-test-gke-tpu-7x -> Failing on develop as well . Known issue
PR-test-gke-g4 -> Test is disabled on develop (reservation not available)
PR-test-gke-h4d -> Reservation not available , test is disabled on develop
PR-test-slurm-gke -> Failing on develop as well with same error
PR-test-gke-tpu-v6e-flex -> Test is disabled. Triggered by babysit
PR-test-gke-managed-lustre -> Failing on develop as well.

@shubpal07

Copy link
Copy Markdown
Contributor Author

Last successful tests on the PR:

Check PR Description / check-description (pull_request)Successful in 4s
cla/google
Dependency Review / dependency-review (pull_request)Successful in 7s
Ensure PR label exists / pr-label-validation (pull_request)Successful in 4s
Label External PRs / label-external (pull_request_target)Successful in 6s
multi-approvers / multi-approvers / multi-approvers (pull_request_review)Successful in 6s
multi-approvers / multi-approvers / multi-approvers (pull_request)Successful in 4s
PR-Go-1-24-build-testSuccessful in 6m
PR-ofe-venv Successful in 1m — Summary
PR-test-gke. Successful in 54m — Summary
PR-test-gke-a2-highgpu-kueue-onspot Successful in 165m — Summary
PR-test-gke-a3-highgpu-onspot Successful in 171m — Summary
PR-test-gke-a3-megagpu Successful in 127m — Summary
PR-test-gke-a3-megagpu-onspot Successful in 108m — Summary
PR-test-gke-a4-onspot Successful in 248m — Summary
PR-test-gke-g4-onspot Successful in 89m — Summary
PR-test-gke-h4d-onspot Successful in 64m — Summary
PR-test-gke-inactive-reservation Successful in 151m — Summary
PR-test-gke-managed-hyperdisk Successful in 45m — Summary
PR-test-gke-storage Successful in 125m — Summary
PR-test-gke-tpu-v6e Successful in 163m — Summary
PR-test-ml-gke Successful in 188m — Summary
PR-test-ml-gke-e2e Successful in 198m — Summary
PR-validation Successful in 8m — Summary
Required
Use pre-commit to validate Pull Request / pre-commit (pull_request)Successful in 15m
Required
Use pre-commit to validate Pull Request / pre-commit-highest-dependencies (pull_request)Successful in 13m

…K storage

Change-Id: I0d12f28ba2eba04d7e9fe4c585ccf9be534069aa
…nd verification jobs

Change-Id: I0c21a7148f7385c329c001bd0cb686905e88310d
…l level and updating readme

Change-Id: I2d6c853bbfd4aba5b58cd1a535a0b31398cf94dd
…al Storage

Change-Id: Ie741ec01d646c8da0450e2d1805f4746de2030fd
…n gke-cluster system node pool

Change-Id: I24ab146763276fccaea0b60a8f5e99b2106b893a
@shubpal07
shubpal07 merged commit e9bfa63 into GoogleCloudPlatform:develop Jun 24, 2026
16 of 82 checks passed
ksaishree pushed a commit to ksaishree/cluster-toolkit that referenced this pull request Jul 1, 2026
ep-nag pushed a commit to nagconsulting/cluster-toolkit that referenced this pull request Aug 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release-key-new-features Added to release notes under the "Key New Features" heading.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants