This tutorial shows you how to fine-tune a Gemma 4 large language model (LLM) on a single-host, multi-GPU Google Kubernetes Engine (GKE) cluster on Google Cloud. This cluster uses a single A4 virtual machine (VM) instance with 8 NVIDIA B200 GPUs attached.
The two main processes described in this tutorial are as follows:
- Deploy a high-performance GKE cluster by using GKE Autopilot. As part of this deployment, you create a custom VM image with the necessary software pre-installed.
- After the cluster is deployed, you run a distributed fine-tuning job by using the set of scripts that accompany this tutorial. The job leverages the Hugging Face Accelerate library.
This tutorial is intended for machine learning (ML) engineers, researchers, platform administrators and operators, and for data and AI specialists who are interested in deploying GKE clusters on Google Cloud to train LLMs.
Objectives
Access the Gemma 4 model by using Hugging Face.
Prepare your environment.
Create and deploy an A4 GKE cluster.
Fine-tune the Gemma 4 model by using the Hugging Face Accelerate library with fully sharded data parallel (FSDP).
Monitor your job.
Clean up.
Costs
In this document, you use the following billable components of Google Cloud:
To generate a cost estimate based on your projected usage,
use the pricing calculator.
Before you begin
Enable the required APIs, if any are not already enabled:
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.gcloud services enable compute.googleapis.com
container.googleapis.com file.googleapis.com logging.googleapis.com cloudresourcemanager.googleapis.com servicenetworking.googleapis.com Enable the default service account for your Google Cloud project:
export PROJECT_NUMBER="$(gcloud projects describe "YOUR_PROJECT_ID" --format "value(project_number)")" gcloud iam service-accounts enable "${PROJECT_NUMBER}-compute@developer.gserviceaccount.com" \ --project=YOUR_PROJECT_ID
Grant the roles that the default service account needs to run the fine-tuning workload:
ROLES=("roles/aiplatform.user" "roles/container.developer" "roles/storage.objectUser") for role in "${ROLES[@]}"; do gcloud projects add-iam-policy-binding YOUR_PROJECT_ID \ --member="serviceAccount:${PROJECT_NUMBER}-compute@developer.gserviceaccount.com" \ --role="${role}" done unset ROLES
Create local authentication credentials for your user account:
gcloud auth application-default login
Enable OS Login for your project:
gcloud compute project-info add-metadata \ --metadata=enable-oslogin=TRUE \ --project=YOUR_PROJECT_ID
To get the permissions that you need to complete this tutorial, ask your administrator to grant you the following IAM roles on your project:
- Kubernetes Engine Admin (
roles/container.admin) - Compute Admin (
roles/compute.admin) - Storage Admin (
roles/storage.admin) - Artifact Registry Administrator (
roles/artifactregistry.admin) - Cloud Build Editor (
roles/cloudbuild.builds.editor) - Service Account User (
roles/iam.serviceAccountUser) - Service Account Admin (
roles/iam.serviceAccountAdmin) - Project IAM Admin (
roles/resourcemanager.projectIamAdmin) - Service Usage Admin (
roles/serviceusage.serviceUsageAdmin)
For more information about granting roles, see Manage access to projects, folders, and organizations.
You might also be able to get the required permissions through custom roles or other predefined roles.
Access Gemma 4 by using Hugging Face
To use Hugging Face to access Gemma 4, do the following:
- Sign in to Hugging Face
- Create a Hugging Face
writeaccess token.
Click Your Profile > Settings > Access tokens > +Create new token - Copy and save the
write accesstoken value. You use it later in this tutorial.
Prepare your environment
To prepare your environment, set the following:
Replace the following:
YOUR_PROJECT_ID: the ID of the Google Cloud project where you want to create the GKE cluster.
YOUR_CLUSTER_NAME: the name of the GKE cluster to create.
YOUR_REGION: the region where you want to create your GKE cluster. You can only create the cluster in the region where your reservation exists.
YOUR_RESERVATION_NAME: the identifier for your reserved capacity.
YOUR_HF_TOKEN: the Hugging Face access token that you created in the previous section.
YOUR_ARTIFACT_REGISTRY_LOCATION: the Google Cloud region where you want to create your Artifact Registry repository. To minimize image pull latency, we recommend that you use the same region that you used for YOUR_REGION.
Create a GKE cluster in Autopilot mode
To create a GKE cluster in Autopilot mode, run the following command:
Creating the GKE cluster might take some time to complete. To verify that Google Cloud has finished creating your cluster, go to Kubernetes clusters on the Google Cloud console.
Create a Kubernetes secret for Hugging Face credentials
To create a Kubernetes secret for Hugging Face credentials, follow these steps:
Configure
kubectlto communicate with your GKE cluster:Create a Kubernetes secret to store your Hugging Face token:
Prepare your workload
To prepare your workload, you do the following:
Create workload scripts
To create the scripts that your fine-tuning workload uses, do the following:
Create a directory for the workload scripts. Use this directory as your working directory.
Create the
cloudbuild.yamlfile to use Google Cloud Build. This file creates your workload container and stores it in Artifact Registry:Create a
Dockerfilefile to execute the fine-tuning job:Create the
accel_fsdp_gemma4_config.yamlfile. This configuration file directs Hugging Face Accelerate to split the tuning job across the eight local GPUs on your single host by using FSDP:Create the
finetune.yamlfile:Create the
finetune.pyfile:
Use Docker and Cloud Build to create a fine-tuning container
Create an Artifact Registry Docker Repository:
In the
llm-finetuning-gemmadirectory that you created in an earlier step, run the following command to create the fine-tuning image and push it to Artifact Registry.Export the image URL. You use it at a later step in this tutorial:
Start your fine-tuning workload
To start your fine-tuning workload, do the following:
Apply the finetune manifest to create the fine-tuning job:
Because you're using clusters in GKE Autopilot mode, it might take a few minutes to start your GPU enabled node.
Monitor the job by running the following command:
After the pods are running, check the logs of the job:
The job resource downloads the model data then fine-tunes the model across all eight of the GPUs. The download takes around five minutes to complete. After the download is complete, the fine-tuning process takes approximately two hours and 30 minutes to complete.
Monitor your workload
You can monitor the use of the GPUs in your GKE cluster to verify that your fine-tuning job is efficiently running. To do so, open the following link in your browser:
When you monitor your workload, you can see the following:
- GPUs usage: for a healthy fine-tuning job, you can expect to see the usage of all of your 8 GPUs rise and stabilize to a high level throughout your training.
- Job duration: the job should take approximately two hours and 30 minutes to complete on the specified A4 cluster.
Clean up
To avoid incurring additional charges, delete the resources created during this tutorial.
Delete your resources
To delete your fine-tuning job, run the following command:
To delete your GKE cluster, run the following command: