There's a beta Compute Engine endpoint that answers "can I get twenty Spot VMs of a given machine type in this region right now?" with a number between zero and one, and most teams on Spot have never called it. A working tour of the capacity-advice APIs (advice.capacity and advice.capacityHistory): the exact request, the exact response, and the roles/compute.viewer plus ADC setup that gets you there. Read the per-zone shards, not just the headline obtainability score. The shards tell you where your capacity actually is. Then how to read the number well: these scores are advisory by design, they move between calls, and a score you read as a reservation will steer you wrong at the worst possible moment. Part 2 of the Obtainability series.
You went to Spot VMs for the discount and did the homework: checkpoint code, a drain handler, a preemption survived. Then a job sat in Pending for days across three regions and never got a single L4. Scarcity grew a second half. Interruption is losing capacity you have; obtainability is failing to get capacity at all, and you can't checkpoint your way out of a stockout. Part 1 of the series that builds the probe the Compute Fallback Ladder left unbuilt.
Every resilience pattern you know assumes the machine shows up. Accelerator scarcity breaks that assumption, and the pattern catalogs have not caught up. Here is the Compute Fallback Ladder—an ordered set of rungs your workload can run on, a selector that picks the highest obtainable one, and a promotion path back up when capacity returns.
Agent density is a lifecycle problem, not an isolation problem. Google's own GKE Agent Sandbox benchmark puts the sandbox at 44% and suspend and resume at up to 3.5x—the sandbox is the smaller half. And upstream Kubernetes 1.37 is quietly turning checkpoint and restore into a kubelet primitive. Here's why the economics live in the lifecycle, not the box.
I'm hiring a Senior Developer Relations Engineer for GKE and AI Infrastructure. Instead of describing the work, I built it: a small, playable cluster scheduler that teaches the real problem behind running AI at scale—gang scheduling, memory ceilings, spot preemption, and a TPU slice with actual topology.