Why accelerator spend runs away

GPU VMs in the A2 and A3 families and Cloud TPU pods bill per accelerator hour at rates far above general compute, so a handful of idle or oversized accelerators moves the bill quickly. Training jobs that hold capacity between runs, and serving endpoints sized for peak that mostly sit warm, are the common culprits.

The discipline that works for compute applies here too, plus one phase specific rule: training and serving have different interruption tolerances, so they should buy capacity differently.

Spot and preemptible for interruptible work

Spot VMs and preemptible accelerators offer deep discounts against on demand in exchange for the provider reclaiming capacity with little notice. Training and batch inference that checkpoint and resume are a natural fit, because an interruption costs minutes, not the job.

Serving traffic that must stay up does not belong on spot. The split is by interruption tolerance: checkpointed training rides spot, live inference rides reserved or committed capacity.

Commitments and reservations for the steady floor

Committed use discounts on GCP lower the rate in exchange for a one or three year commitment, and they suit the steady serving floor that runs continuously. Reservations hold accelerator capacity so a critical endpoint is not starved at peak, separate from the discount decision.

Commit only the floor that runs regardless of demand, the same risk adjusted rule as any commitment, because an accelerator commitment bills whether or not the GPUs are busy. Right sizing comes first: pick the smallest accelerator type and count that meets the latency and throughput target before committing to anything.

A worked example

Indicative figures, verified against the client's billing data, anonymized. A scaling fintech ran training and serving on a single always on GPU pool.

Indicative monthly accelerator spend before and after the split
WorkloadBeforeAfter
Training on always on on demand GPUs60,000 USD22,000 USD on spot
Serving at peak size, mostly idle34,000 USD18,000 USD right sized
No commitment on steady serving floorn/afloor on CUD, ~15 percent lower rate
Total monthly accelerator94,000 USD40,000 USD

Your next step

Split training from serving by interruption tolerance, right size the accelerators, then commit only the steady floor. For the full method read the GCP cost optimization guide, and for neighbouring detail see spot VMs and preemptible economics on GCP and commitment coverage targets on GCP. To apply it, our GCP cost optimization service turns the plan into verified savings, and you can request a free trial.

Independent · buyer-side

Put a defensible number on your cloud spend.

No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.

Buyer-side intelligence, monthly.

The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.