AI NewsInfrastructureAnnouncement
Google says GKE Pod snapshots load a 70-billion-parameter model in 37 seconds, cutting startup latency by up to 89 percent
Google published benchmark results for GKE Pod snapshots on 21 September, saying an 8-billion-parameter model loads in 15 seconds and a 70-billion-parameter model in 37 seconds, cutting cold-start latency by as much as 89 percent by restoring the pod's saved CPU and GPU memory instead of running the load path.

Image: InfoQ
Why it mattersA team can now start a large model server on demand and shut it down when the job is done, instead of paying for a warm pool, so scaling large-model inference to real user traffic becomes an operations question about snapshot lifetimes rather than about how long the model takes to load.
Loading a 70-billion-parameter model into a new server has been the wall that keeps teams paying for a warm pool: a fresh replica takes minutes to be useful, and traffic will not wait. Google Cloud published a blog post on 21 September titled GKE Pod snapshots, reporting that the same replica now comes back in 37 seconds by restoring the pod's saved CPU and GPU memory instead of running the load path.
The feature is checkpoint and restore. Google says the snapshot holds the running application's file descriptors, threads, CPU registers, memory, the container root filesystem, and EmptyDir and tmpfs volumes. A new replica picks up from there and skips the initialization that loads the model. Google's headline numbers: up to 89 percent lower startup latency, an 8-billion-parameter model loading in 15 seconds, and a 70-billion-parameter model loading in 37 seconds. The feature reached general availability in May on Google Kubernetes Engine 1.35.3-gke.1234000 and later.
Google's customer example
Codeway's Retake platform had built its own caching layer for compiled artifacts and had cold starts down to about a minute. Its lead DevOps engineer, Ahmet Furkan Comak, is quoted in the Google Cloud post saying Pod snapshots cut that to "just 8 seconds", and that the team now starts H100 instances for a specific job and shuts them down when the job is finished. Google published no independent measurement of that figure, so treat the 8 seconds as Codeway's own statement about their own workload.
What InfoQ's reporting adds
InfoQ covered the release on 26 September and framed the harder question as what happens after a snapshot exists. The story quotes Mohana Narasimha G., a senior DevOps and MLOps engineer, on LinkedIn: "The restore path is compelling, but I suspect snapshot invalidation will be the harder platform problem than capture itself. Model digest, CUDA/driver version, GPU type/topology, and runtime config all become part of the compatibility key."
InfoQ names the practical costs of that compatibility key, quoting the Google documentation: whole-pod snapshots need matching gVisor kernel and GPU driver versions and the same machine series, so upgrading a node pool can invalidate every snapshot silently. When no compatible snapshot exists, the pod starts normally with no error, and the fast-start benefit goes away. Restore is also gradual: gVisor's kernel returns first, the application runs at that point, and memory keeps loading in the background. E2 machine types are out entirely, multi-GPU pods are supported only on L4 GPUs, and Multi-Instance GPU sharing is unsupported.
A snapshot in Cloud Storage holds the complete memory of a running workload, which in an agent-sandbox setup is memory that has executed model-generated code. Access rests on Workload Identity Federation and per-pod IAM bindings, and Google notes those bindings can take time to propagate.
The takeaway for a team building with models is about staffing. A model server that takes a minute to warm up gets served from a warm pool at all hours, because a minute of user waiting is a minute the app is broken. A model server that starts in 37 seconds turns the operations question from "how many spare replicas do I always keep" into "how long does a snapshot stay valid before an upgrade breaks it". The second question has a much cheaper answer.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


