GCP sells the same compute under several pricing models, and the model you choose moves the rate far more than incremental rightsizing does. On demand is the default and the most expensive per unit. Sustained use discounts apply automatically and for free when an eligible Compute Engine instance runs for most of the month. Committed use discounts trade a one or three year commitment for a much larger reduction, up to around 57 percent for three year resource based commitments, which is indicative of current pricing. Spot virtual machines give the deepest discount for interruptible work. And some services, BigQuery being the clearest, offer a separate capacity pricing model that replaces per usage charges with reserved throughput. The buyer skill is matching each workload to the model whose flexibility tradeoff it can actually accept.
Get the model wrong and you either overpay on demand for steady load or strand a commitment on load that moved. Here is each model, what it costs you in flexibility, and the workload it belongs on.
On demand and sustained use: the no commitment baseline
On demand is the list rate with no commitment and full flexibility to start and stop. Sustained use discounts sit on top of it for eligible Compute Engine instances and apply automatically as an instance accumulates running hours across the month, with the discount growing as usage rises toward full month running. You do nothing to earn them and carry no commitment, which makes them the right floor for genuinely variable compute that you cannot forecast. The catch is that they are modest compared with a commitment, so leaving steady baseline load on sustained use alone means paying far more than that same load would cost under a committed use discount.
Committed use discounts: the biggest lever and the biggest risk
Committed use discounts, or CUDs, are where the real money is. You commit to a one or three year level of usage and receive a steep discount, and they come in two forms that behave differently. Resource based CUDs commit to a specific machine type and region for the largest discounts but the least flexibility. Spend based CUDs commit to an hourly dollar amount on a service family for a smaller discount with far more flexibility about what fills it. The risk is real and it is yours: a CUD bills whether or not you use it, so an aggressive commitment on load that later shrinks or migrates becomes stranded spend. Coverage should follow a defensible forecast of the baseline that will still be there in a year, not the discount you wish you could capture.
Which model does each workload belong on?
The table maps common workload shapes to the model that fits, with the tradeoff each one accepts.
| Workload shape | Best fit model | Why |
|---|---|---|
| Steady baseline compute that runs all year | Committed use discount, one or three year | Largest discount on load you can forecast and will keep |
| Variable compute you cannot forecast | On demand with automatic sustained use discounts | No commitment, modest automatic saving, full flexibility |
| Fault tolerant batch and stateless workers | Spot virtual machines | Deepest discount in exchange for possible preemption |
| Predictable heavy analytics | BigQuery capacity pricing (reserved slots) | Fixed throughput cost beats per data processed at scale |
| Spiky or light analytics | BigQuery on demand by data processed | Pay nothing between queries, no idle slot cost |
Most estates run a blend: a committed use floor under the steady baseline, sustained use and on demand for the variable layer on top, Spot for batch, and the right BigQuery model for the analytics pattern. The mistake is applying one model to everything.
A worked example of model fit
A Fortune 500 retailer ran its entire GCP fleet on demand, taking only the automatic sustained use discounts. Analysis of a representative quarter showed a stable baseline that had not dropped in a year sitting under spiky daytime demand, plus a nightly batch tier and a heavy BigQuery reporting layer billed on demand. We placed three year resource based CUDs under the stable baseline, left the spiky layer on sustained use, moved the batch tier to Spot, and switched the reporting layer to BigQuery capacity pricing sized to its real slot demand. Blended compute and analytics cost fell sharply with no architecture rewrite, and the figures are verified against billing data and anonymized.
Pull a quarter of usage and split it into the load that never went away and the load that came and went. The first belongs on a commitment, the second does not. If everything is on demand, you are overpaying on the steady half.
Where this fits in the GCP estate
Understanding the models is the start; reading them on your own bill is the next step. See how the automatic discount works in decoding GCP credits and how to use them, learn to find these line items in reading your GCP bill line by line, and avoid the traps in common GCP billing surprises. The full picture lives in the GCP cost optimization guide, which links across to the cross cloud cost optimization guide.
Frequently asked questions
What pricing models does GCP offer?
What is the difference between sustained use and committed use discounts?
Which GCP pricing model is cheapest?
Get the buyer side GCP pricing playbook
We match every GCP workload to the pricing model it belongs on and size commitments to a forecast you can defend, so you capture the discount without stranding spend. Our guarantee: we reduce your cloud spend or we reimburse our service fee. Pricing is either a Fixed Fee scoped up front or Gainshare, a share of verified savings with no retainer and no risk.
Put a defensible number on your cloud spend.
No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.
The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.