Skip to content

· 11 min read · Engineering

Obtainability, Part 2: The API That Tells You What You Can Get

Obtainability, Part 2: The API That Tells You What You Can Get

Part 1 ended on a teaser about a beta Compute Engine endpoint that will tell you, right now, whether you can get twenty Spot VMs of a given machine type in a given region, and it returns a field called obtainability. Most folks don’t know about it yet, and I’m here to change that.

I’ll show you the exact request, the exact response, and the two-line IAM setup that gets you there. Then I’ll show you what the number can tell you and what it can’t. An obtainability score you treat as a reservation will pick the worst possible moment to be wrong. It did that to me across three regions, and that story is Part 1. This post is the tool I built afterward.

This series has a companion GitHub repo, cwest/gke-spot-playbook. One tool in it, capacity-advisor, is a small open-source Go program I wrote that calls these two endpoints and prints a ranked table. You give it a config, a set of candidate machine types and the regions you’d run in, and it queries the capacity-advice APIs for each combination, then ranks where your Spot capacity actually is. It lives in advisor/ in that repo, and every request and response you’ll see below comes from running it. It’s my tool, not Google Cloud’s console feature of a similar name.

The Request: Ask a Region a Question

The endpoint is advice.capacity, one of the Compute Engine capacity-advice beta APIs. You POST to a region and describe the fleet you wish you had:

POST https://compute.googleapis.com/compute/beta/projects/PROJECT/regions/us-central1/advice/capacity

The body is a fleet request, not a single-VM request. You’re not asking “does this one machine exist”; you’re asking “can this region give me twenty of these, spread however it likes, right now.”

{
"instanceProperties": {
"scheduling": { "provisioningModel": "SPOT" }
},
"instanceFlexibilityPolicy": {
"instanceSelections": {
"selection-1": {
"machineTypes": ["n2-standard-2", "n2-standard-4"]
}
}
},
"size": 100,
"distributionPolicy": { "targetShape": "ANY" }
}

provisioningModel: SPOT says score me for Spot, not on-demand. This is a Spot-capacity question; the whole reason obtainability is interesting is that Spot is where capacity actually goes missing. It’a also where capacity can be found.

instanceSelections is a map, and machineTypes inside it is a list. You hand the API a set of machine types you’d accept, and the region tells you which of them it can actually fill. It’s a “here’s what I’d take” list, not a “here’s the one thing I want.” The API caps that list at five: advice.capacity accepts at most five machine types per request, so capacity-advisor sends them in batches of five and my configs stay under that in the first place. That’s a comfortable ceiling, because most of the time there’s an ideal machine type and a few acceptable tradeoffs. Give it the machine types you’d genuinely run, in the order you’d prefer them, and no more.

size is how many VMs you want. The API scores against that target, so size: 100 and size: 4 can come back with different obtainability for the same machine type in the same region. Scarcity isn’t a property of a zone. It’s a property of a zone being asked for a specific amount.

targetShape: ANY tells the region to place the fleet wherever it can, packing capacity across zones rather than forcing it all into one. The API also accepts ANY_SINGLE_ZONE and BALANCED; ANY is the one you want when the question is “can I get this at all,” because it gives the region the most freedom to say yes.

Send that, and a region answers. But first you need permission to ask.

The Setup: Two Lines and a Read-Only Role

You don’t need write access to anything to ask this question. You’re reading a score, not creating a VM, so roles/compute.viewer on the project is enough.

Authenticate with Application Default Credentials and you’re done:

Terminal window
gcloud auth application-default login

Enable the Compute Engine API on the project, have the roles/compute.viewer role, and run that one command.

The Response: The Shards Beat the Score

Here’s what comes back for the request above:

{
"recommendations": [
{
"scores": { "obtainability": 0.9, "estimatedUptime": "600s" },
"shards": [
{
"instanceCount": 90,
"machineType": "n2-standard-2",
"zone": ".../zones/us-central1-a",
"provisioningModel": "SPOT"
},
{
"instanceCount": 10,
"machineType": "n2-standard-4",
"zone": ".../zones/us-central1-c",
"provisioningModel": "SPOT"
}
]
}
]
}

The headline is obtainability: 0.9. Google’s own reference defines it as “the likelihood of successfully obtaining (provisioning) the requested number of VMs,” on a scale from zero to one, higher being better. So 0.9 reads as “you’ll probably get your hundred VMs.” Useful. But it’s not the whole picture and the details matter.

Read the shards instead. The response didn’t just score the region; it told you how the region would fill the order. Ninety n2-standard-2 in us-central1-a, ten n2-standard-4 in us-central1-c. That’s a placement plan. It tells you which zones actually have your capacity and which of your accepted machine types the region is leaning on. If ninety of your hundred VMs land in one zone. If you were after “multi-zone” resilience this response isn’t it. Thankfully, for spot-friendly workloads, this should matter less.

The other score, estimatedUptime, is the expected run time before preemption for the majority of your Spot VMs, expressed as a duration string like "600s". It’s best-effort and drawn from historical data, and it’s the answer to a different question than obtainability. Not “can I get the capacity” but “how long will it stay.” A workload that checkpoints every ten minutes cares about a ten-minute uptime estimate very differently than one that needs an uninterrupted hour.

The History Endpoint: What the Zone Has Been Doing

The second endpoint, advice.capacityHistory, answers a question the point-in-time score can’t. Not “can I get this now” but “how has this zone behaved lately.” You pin it to a machine type and, optionally, a single zone:

{
"instanceProperties": {
"machineType": "n2-standard-32",
"scheduling": { "provisioningModel": "SPOT" }
},
"locationPolicy": { "location": "zones/us-central1-a" },
"types": ["PREEMPTION", "PRICE"]
}

One detail here cost me a debugging session, so I’ll save you the trouble. The locationPolicy block is how you scope to a zone, and to ask for region-wide history you omit the block entirely rather than sending it empty. An empty location isn’t “everywhere”; it’s a malformed request. Include the block for a zone, drop the whole block for a region.

The response carries two series. preemptionHistory is a list of daily records, each with an interval and a preemptionRate between zero and one, defined as the fraction of Spot VMs that were preempted (preempted Spots over Spots that stopped running). priceHistory carries a listPrice per interval, given as units and nanos you reassemble into a dollar figure.

{
"preemptionHistory": [
{ "interval": { "startTime": "2026-04-20T07:00:00Z" }, "preemptionRate": 0.52 },
{ "interval": { "startTime": "2026-04-21T07:00:00Z" }, "preemptionRate": 0.31 }
],
"priceHistory": [
{ "interval": { "startTime": "2026-04-27T07:00:00Z" }, "listPrice": { "currencyCode": "USD", "nanos": 478720000 } }
]
}

A nanos of 478,720,000 is $0.47872 an hour. A preemption rate that reads 0.52 then 0.31 is a zone whose Spot pool churned through half its instances one day and a third the next. Neither number is an obtainability score, and that’s exactly why they’re worth having. Obtainability guesses at whether you can get in; the history tells you what happened to the people who did.

This is important: the API makes no promise about the order of these records, so don’t trust array position. Sort the preemption records by their interval start before you read them as a series, and pick the price with the latest interval as “current” rather than grabbing the last element.

How to Read the Number

Everything above is documented behavior. What follows isn’t in the reference. It’s how to read the number well, which is the part I had to learn through trial and error.

These scores are advisory in the reference and in the word “beta” on the endpoint. It’s easy to read that sentence, ignore its meaning, and then go wire the number into a scheduler as if it were inventory. I did. My job was small: one LoRA fine-tune of a Gemma model on a single L4 GPU, one accelerator for one workload. capacity-advisor called the endpoint and read back obtainability: 0.90 for us-east1-c, so I set that zone as the target and let the autoscaler chase it. The autoscaler came back empty, over and over, for about thirty-six hours across three regions, while the score kept reading high. The score was doing exactly its job the whole time. It was reporting a probability, and I was reading it as a promise.

Three things I learned by leaning on the number too hard:

A score is a probability, and probabilities have a shelf life. Obtainability is computed from historical data and current conditions, and “current” is measured in minutes for scarce machine types. I watched one region’s score fall from 0.90 to 0.10 in about four hours. Both readings were accurate for the moment they were taken. If your loop reads the score once and caches it, you’re operating from memory, and memory drifts from reality fast.

A high score answers a narrower question than “is the capacity mine.” There’s a category difference between a prediction that a zone is probably fine and a confirmation that a machine was just allocated to you. Obtainability is the first. Only an actual provisioning attempt is the second, and for genuinely scarce hardware the gap between them is where your outage lives. The number is honest about what it measures. Read it as an answer to “how likely,” not “is it done.”

The estimate covers the request you sent, not the one you’ll send next. Change size, change the machine-type list, change the region, and you’ve asked a different question. A 0.9 for a hundred small VMs tells you nothing reliable about four large ones in the same region. The request you sent is the only request that got scored. If you’re asking a different capacity question, re-score before you act.

flowchart TD
    q["advice.capacity: score the fleet"] --> s{"obtainability high?"}
    s -->|"no"| skip["Don't bother: try another rung"]
    s -->|"yes"| shards["Read the shards: which zones, how much"]
    shards --> attempt["Attempt to provision"]
    attempt -->|"got it"| run["Run the workload"]
    attempt -->|"refused"| learn["Record the refusal: re-score, the reading has moved"]
    learn --> q

Read this way, the API earns its place. It’s a ranking signal, which is a real and valuable thing to have when you’re choosing among regions and machine types you’d otherwise pick by superstition. A 0.9 zone genuinely is a better first attempt than a 0.3 zone. Read the score as what it is, the best available guess about where to try first, and then confirm it the only way capacity is ever confirmed. Ask the question, read the shards, then go find out for real.

What You Can Do With This Today

You can run the request in this post against your own project right now with nothing but a viewer role, and you should. If you want the full working version, it’s the capacity-advisor in the companion repo. That client caps the machine-type list, sorts the history correctly, reassembles the price from units and nanos, and turns the whole thing into a ranked table:

Read the primary source too. The full request and response schemas, every enum, and the exact field definitions live in Google’s reference for advice.capacity and advice.capacityHistory. It’s a beta API, so read the current version. Beta field names move, and this post is a snapshot.

Part 1 said the Compute Fallback Ladder named a probe() and never built it. This score is the closest thing Google ships to that probe, and Part 2’s whole argument is that it’s close but not quite. It’s a prior, not a probe. Part 3 takes the four signals this API surfaces, obtainability and uptime and preemption and price, and turns them into a ranking a scheduler can act on, which is where the first real judgment enters. Because a number is not a node pool, and the distance between them is the rest of this series.