TL
The short answer

GCP sells the same compute under several pricing models, and the model you choose moves the rate far more than incremental rightsizing does. On demand is the default and the most expensive per unit. Sustained use discounts apply automatically and for free when an eligible Compute Engine instance runs for most of the month. Committed use discounts trade a one or three year commitment for a much larger reduction, up to around 57 percent for three year resource based commitments, which is indicative of current pricing. Spot virtual machines give the deepest discount for interruptible work. And some services, BigQuery being the clearest, offer a separate capacity pricing model that replaces per usage charges with reserved throughput. The buyer skill is matching each workload to the model whose flexibility tradeoff it can actually accept.

Get the model wrong and you either overpay on demand for steady load or strand a commitment on load that moved. Here is each model, what it costs you in flexibility, and the workload it belongs on.

On demand and sustained use: the no commitment baseline

On demand is the list rate with no commitment and full flexibility to start and stop. Sustained use discounts sit on top of it for eligible Compute Engine instances and apply automatically as an instance accumulates running hours across the month, with the discount growing as usage rises toward full month running. You do nothing to earn them and carry no commitment, which makes them the right floor for genuinely variable compute that you cannot forecast. The catch is that they are modest compared with a commitment, so leaving steady baseline load on sustained use alone means paying far more than that same load would cost under a committed use discount.

Committed use discounts: the biggest lever and the biggest risk

Committed use discounts, or CUDs, are where the real money is. You commit to a one or three year level of usage and receive a steep discount, and they come in two forms that behave differently. Resource based CUDs commit to a specific machine type and region for the largest discounts but the least flexibility. Spend based CUDs commit to an hourly dollar amount on a service family for a smaller discount with far more flexibility about what fills it. The risk is real and it is yours: a CUD bills whether or not you use it, so an aggressive commitment on load that later shrinks or migrates becomes stranded spend. Coverage should follow a defensible forecast of the baseline that will still be there in a year, not the discount you wish you could capture.

Which model does each workload belong on?

The table maps common workload shapes to the model that fits, with the tradeoff each one accepts.

Workload shapeBest fit modelWhy
Steady baseline compute that runs all yearCommitted use discount, one or three yearLargest discount on load you can forecast and will keep
Variable compute you cannot forecastOn demand with automatic sustained use discountsNo commitment, modest automatic saving, full flexibility
Fault tolerant batch and stateless workersSpot virtual machinesDeepest discount in exchange for possible preemption
Predictable heavy analyticsBigQuery capacity pricing (reserved slots)Fixed throughput cost beats per data processed at scale
Spiky or light analyticsBigQuery on demand by data processedPay nothing between queries, no idle slot cost

Most estates run a blend: a committed use floor under the steady baseline, sustained use and on demand for the variable layer on top, Spot for batch, and the right BigQuery model for the analytics pattern. The mistake is applying one model to everything.

A worked example of model fit

Worked example

A Fortune 500 retailer ran its entire GCP fleet on demand, taking only the automatic sustained use discounts. Analysis of a representative quarter showed a stable baseline that had not dropped in a year sitting under spiky daytime demand, plus a nightly batch tier and a heavy BigQuery reporting layer billed on demand. We placed three year resource based CUDs under the stable baseline, left the spiky layer on sustained use, moved the batch tier to Spot, and switched the reporting layer to BigQuery capacity pricing sized to its real slot demand. Blended compute and analytics cost fell sharply with no architecture rewrite, and the figures are verified against billing data and anonymized.

The buyer test

Pull a quarter of usage and split it into the load that never went away and the load that came and went. The first belongs on a commitment, the second does not. If everything is on demand, you are overpaying on the steady half.

Where this fits in the GCP estate

Understanding the models is the start; reading them on your own bill is the next step. See how the automatic discount works in decoding GCP credits and how to use them, learn to find these line items in reading your GCP bill line by line, and avoid the traps in common GCP billing surprises. The full picture lives in the GCP cost optimization guide, which links across to the cross cloud cost optimization guide.

Frequently asked questions

What pricing models does GCP offer?
GCP bills compute on demand by default, applies sustained use discounts automatically when a virtual machine runs for a large share of the month, offers committed use discounts in exchange for a one or three year commitment, and prices some services such as BigQuery either on demand by data processed or on capacity by reserved slots. Each model trades flexibility against a lower rate.
What is the difference between sustained use and committed use discounts?
Sustained use discounts apply automatically with no commitment when eligible Compute Engine instances run for a large portion of the month, giving a modest discount for doing nothing. Committed use discounts require you to commit to a one or three year level of usage or spend and give a much larger discount, typically up to around 57 percent for three year resource based commitments, in exchange for carrying utilization risk.
Which GCP pricing model is cheapest?
There is no single cheapest model, only the cheapest fit for a workload. Steady baseline compute is cheapest on committed use discounts, variable compute that cannot be committed is cheapest left to take automatic sustained use discounts, fault tolerant batch is cheapest on Spot, and predictable heavy BigQuery is cheapest on capacity pricing rather than on demand.

Get the buyer side GCP pricing playbook

We match every GCP workload to the pricing model it belongs on and size commitments to a forecast you can defend, so you capture the discount without stranding spend. Our guarantee: we reduce your cloud spend or we reimburse our service fee. Pricing is either a Fixed Fee scoped up front or Gainshare, a share of verified savings with no retainer and no risk.

Independent · buyer-side

Put a defensible number on your cloud spend.

No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.

Buyer-side intelligence, monthly.

The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.