Sitemap
Google Cloud - Community

A collection of technical articles and blogs published or curated by Google Cloud Developer Advocates. The views expressed are those of the authors and don't necessarily reflect those of Google.

Still Packaging AI Models in Containers? Do This Instead on Cloud Run

--

If you’ve ever waited for a massive Docker image to build and push just because you changed a single line of application code, you know the pain of traditional LLM model deployment. The process of baking large model files directly into container images is a common practice, but it’s one that leads to slow development cycles, bloated artifacts, and frustrating deployment limits.

There’s a more elegant solution. By decoupling your model files from your application code, you can create a far more efficient and scalable serving architecture on Google Cloud. This post will walk you through a practical, step-by-step guide to achieving this using Cloud Run, Cloud Storage FUSE, and Ollama.

The Architecture: Decoupling Models with Cloud Storage FUSE

This approach treats your model as a distinct piece of data, not as a part of your application’s build artifact. Instead of copying the model into your Docker image, we’ll store it in a Google Cloud Storage bucket.

This is where Cloud Storage FUSE comes in. FUSE (Filesystem in Userspace) is a technology that allows you to create a file system in user space. The Cloud Storage FUSE adapter, specifically, lets you mount a Cloud Storage bucket so that it appears as a local directory inside your Cloud Run container. Your application can read the model files as if they were on the local disk, but they are actually being streamed on-demand from Cloud Storage.

This simple architectural change unlocks several benefits:

  • No more image size limits: Your container image remains small and nimble, containing only your application code. The model, regardless of its size, lives in Cloud Storage.
  • Rapid development: When you update your application, you’re only deploying a small container image, reducing deployment times from minutes (or hours) to seconds.
  • Centralized model management: You can store a single model artifact in Cloud Storage and have multiple Cloud Run services or jobs mount it, ensuring consistency and saving on storage costs.

A Note on Network Performance

To ensure your models load quickly and efficiently, you need to optimize the network path between your Cloud Run service and your Cloud Storage bucket. Google’s best practices recommend the following configuration:

  • Enable Direct VPC Egress: You must configure your Cloud Run service to send all outbound traffic through your Virtual Private Cloud (VPC) network. This is done by setting the VPC egress setting to all-traffic.
  • Enable Private Google Access: With Direct VPC egress enabled, you also need to ensure that Private Google Access is enabled for the subnet your traffic flows through. This allows your Cloud Run service to reach Google APIs and services, like Cloud Storage, using Google’s internal network, which is significantly faster and more reliable.

For a detailed comparison of different model loading strategies and their trade-offs, refer to the official Cloud Run documentation on best practices for AI inference. This guide provides a valuable table that breaks down the implications of each approach on deployment time, startup time, and cost.

The Workflow: A Two-Part Deployment

Our deployment process is divided into two logical parts: a one-time job to stage the model and a long-running service to serve it.

Part 1: Transferring the Model to Cloud Storage with a Cloud Run Job

Press enter or click to view image in full size

First, we need to get our model from its source — in this case, the Hugging Face Hub — and place it into our Cloud Storage bucket. This is a perfect use case for a Cloud Run Job, which is designed for run-to-completion tasks.

The job leverages obstore, a high-performance, asynchronous Python library designed for exactly this type of task: streaming data from one object store (like the Hugging Face Hub) to another (like Cloud Storage). While standard tools like hf_hub_download are excellent for downloading files to a local machine, they would require a two-step process in our job: first downloading the model to the container’s temporary disk, and then uploading it to Cloud Storage. obstore allows us to bypass the intermediate step by streaming the file directly, which is far more efficient in a containerized environment. It handles the concurrent, in-memory streaming automatically, which simplifies the code significantly.

To accomplish this, we’ll package our transfer logic into a Python script. This script will be the heart of our Cloud Run Job, and looks like this:

# Initialize stores for the source (Hugging Face) and destination (Cloud Storage)
http_store = HTTPStore.from_url("https://huggingface.co", client_options=client_options)
gcs_store = GCSStore(bucket=config["gcs_bucket_name"])

# Get a streaming response from the source
streaming_response = await obs.get_async(http_store, download_path)

# Stream the response directly to the destination
await obs.put_async(gcs_store, gcs_destination_path, progress_stream)

To handle authentication securely, the Hugging Face API token is stored in Google Secret Manager and mounted into the Cloud Run Job as a file. This avoids exposing sensitive credentials in environment variables or the container image.

Here is the command to deploy the job:

gcloud beta run jobs deploy your-job-name \\
--image your-job-image \\
--region us-central1 \\
--cpu 2 \\
--memory 4Gi \\
--service-account your-service-account \\
--labels dev-tutorial=blog-gcsfuse \\
--set-env-vars HF_REPO_ID=unsloth/gemma-3n-E4B-it-GGUF,HF_MODEL_FILE_PATTERN=*Q4_K_XL*,GCS_BUCKET_NAME=your-bucket-name,GCS_MODEL_PATH_PREFIX=unsloth-gemma-3n-e4b-it-gguf-model/ \\
--set-secrets=/etc/secrets/hf-token/HF_TOKEN=HF_TOKEN:latest \\
--project YOUR_PROJECT_ID \\
--task-timeout 24h \\
--execute-now

Part 2: Serving the Model with a Cloud Run Service

Press enter or click to view image in full size

With our model safely stored in Cloud Storage, we can now deploy our inference server. We’ll use Ollama, a popular and easy-to-use tool for serving large language models.

The magic happens during the deployment of our Cloud Run service. We use two specific flags in our gcloud run deploy command:

gcloud run deploy your-service-name \\
--source ./ollama_service \\
--region us-central1 \\
--allow-unauthenticated \\
--project YOUR_PROJECT_ID \\
--labels dev-tutorial=blog-gcsfuse \\
--gpu 1 \\
--gpu-type nvidia-l4 \\
--add-volume=name=ollama-gcs-models,type=cloud-storage,bucket=YOUR_BUCKET_NAME,readonly=true \\
--add-volume-mount=volume=ollama-gcs-models,mount-path=/models \\
--add-volume=name=ollama-writable-state,type=in-memory,size-limit=1Gi \\
--add-volume-mount=volume=ollama-writable-state,mount-path=/var/lib/ollama

The --add-volume and --add-volume-mount flags are the linchpin of this architecture. They tell Cloud Run to mount our Cloud Storage bucket at the /models path inside our container.

Now, the Ollama container can see the model files at /models. However, Ollama uses a content-addressable storage system, meaning it identifies models by the SHA256 hash of their content, not by their filename. To bridge this gap, we use a simple entrypoint script. When the container starts, this script:

  1. Calculates the SHA256 hash of the model file located at /models/….
  2. Creates a symbolic link from the model file to the location where Ollama expects to find its model blobs (/var/lib/ollama/blobs/sha256-THE_HASH).
  3. Uses the ollama create command with a simple Modelfile that points to the model path.
  4. Starts the Ollama server.

Here’s the condensed script, showing just the essential commands:

#!/bin/sh
set -e

MODEL_FILE_PATH="/models/unsloth-gemma-3n-e4b-it-gguf-model/gemma-3n-E4B-it-UD-Q4_K_XL.gguf"
MODEL_NAME="gemma-3n-custom"
BLOBS_DIR="/var/lib/ollama/blobs"

ollama serve &
sleep 3

MODEL_SHA256=$(sha256sum "$MODEL_FILE_PATH" | awk '{print $1}')
BLOB_PATH="$BLOBS_DIR/sha256-$MODEL_SHA256"

mkdir -p "$BLOBS_DIR"
ln -s "$MODEL_FILE_PATH" "$BLOB_PATH"

ollama create "$MODEL_NAME" -f /workspace/Modelfile

wait $!

This clever use of a symbolic link enables Ollama to use the file directly from the Cloud Storage FUSE mount without needing to copy it.

A Better Path Forward

By separating the concerns of our application code and our model artifacts, we can build more robust, scalable, and developer-friendly AI systems. This Cloud Run architecture provides a clear path to escape the limitations of monolithic container images and embrace a more dynamic and efficient way of serving large language models.

To get started and see the full implementation, check out the complete code and step-by-step instructions in the notebook. For a deeper dive into the technologies, the official documentation for Cloud Run, Cloud Storage FUSE, and Ollama are excellent resources.

I’d love to hear what you build with this architecture. Connect with me on LinkedIn or X and share your projects!

--

--

Karl Weinmeister
Karl Weinmeister

Written by Karl Weinmeister

Developer Relations @ Google Cloud

Google Cloud - Community
Google Cloud - Community

Published in Google Cloud - Community

A collection of technical articles and blogs published or curated by Google Cloud Developer Advocates. The views expressed are those of the authors and don't necessarily reflect those of Google.