TL
The short answer

On demand inference charges per token consumed and provisioned throughput charges a fixed rate for reserved model capacity, so on demand is cheaper for variable or low volume traffic and provisioned throughput is cheaper once sustained demand keeps the reserved capacity highly utilized. Provisioned throughput also guarantees latency and capacity that on demand endpoints can throttle under load, so high volume production traffic often chooses it for predictability as well as price.

Here is how each model bills, where the crossover sits, what provisioned throughput buys beyond price, and a worked example of the decision for a real traffic pattern.

How does each pricing model bill?

On demand inference, the default on every managed model platform, charges per thousand input and output tokens with no commitment. You pay only for what you send, and the rate per token is the same whether you send one request a day or thousands a second, until you hit a throughput quota and requests start to throttle. Provisioned throughput instead reserves a fixed block of model capacity, billed by the hour for a monthly or longer term, and within that block you can run as many tokens as the capacity supports at no extra per token charge. AWS Bedrock sells this as Provisioned Throughput, Azure OpenAI as Provisioned Throughput Units, and Google Vertex AI as provisioned throughput, each with monthly or annual commitments that lower the effective hourly rate.

Where is the crossover?

The crossover is the volume at which the fixed reserved rate equals what you would pay per token on demand for the same work. Below it, on demand is cheaper because you are not paying for idle reserved capacity; above it, provisioned throughput is cheaper because the reserved block absorbs more tokens at no marginal cost. The reserved capacity only pays off if you keep it busy, so the real question is your sustained utilization across the day, not your peak. Traffic that runs hot for two hours and quiet for twenty two rarely justifies a full day of reserved capacity; traffic that holds a high floor around the clock crosses over quickly.

Traffic patternCheaper modelWhy
Low or unpredictable volumeOn demandNo idle reserved capacity to pay for
Spiky, with long quiet periodsOn demand, or a small reserved floorReserved capacity sits idle between spikes
High and steady around the clockProvisioned throughputReserved block stays busy, no per token charge above it
High floor plus spikesReserved floor plus on demand burstCommit the floor, burst the rest per token

What does provisioned throughput buy beyond price?

Cost is only half the decision. Provisioned throughput reserves capacity, so it delivers predictable latency and a guaranteed request rate that on demand endpoints cannot promise once a region is busy. For interactive products where a slow or throttled response costs conversions, that guarantee can justify reserved capacity even slightly before the pure price crossover. The mirror risk is utilization: reserved capacity you do not use is pure waste, the same use it or lose it shape as a compute commitment, so size the reservation to a defensible forecast of sustained demand rather than to the peak you hope to reach.

A worked example

Worked example

A software company ran a customer facing assistant entirely on demand. Daytime traffic held a high, steady floor for roughly ten hours, then fell away overnight. Modelling the floor showed that reserving capacity to cover the steady daytime band, while leaving the overnight tail and occasional spikes on demand, cost less than paying per token for the whole floor, and removed the throttling that had been hurting response times at peak. Committing the full twenty four hours would have stranded capacity overnight and cost more. The decision was to reserve the predictable band and burst the rest. Figures are verified against billing data and anonymised; rates are indicative and depend on model, region, and term.

Frequently asked questions

Is provisioned throughput cheaper than on demand inference?
Only above the crossover volume, where your sustained traffic keeps the reserved capacity busy. Below it you pay for idle capacity and on demand is cheaper. Model your sustained tokens per hour against the reserved rate before committing.
When should I reserve inference capacity?
When traffic holds a high, predictable floor for many hours a day, or when you need guaranteed latency and request rates that on demand endpoints throttle under load. Reserve the defensible floor and leave spikes and quiet periods on demand.
What is the risk of provisioned throughput?
Underutilization. Reserved capacity is a use it or lose it commitment, so capacity you do not consume is wasted spend. Size the reservation to a forecast you can defend, not to an optimistic peak.

Right size your inference commitment

We model on demand against provisioned throughput across Bedrock, Azure OpenAI, and Vertex, find your crossover, and size reservations to a forecast you can defend, as an independent advisory that takes zero provider commissions and answers only to you. Our guarantee: we reduce your cloud spend or we reimburse our service fee, on a Fixed Fee or a no risk Gainshare basis. Download the cloud cost optimization playbook, read the cross cloud cost optimization guide, and on the unit math see token economics for enterprise buyers. For monthly buyer side analysis, subscribe to The Cloud Spend Navigator.

Independent · buyer-side

Put a defensible number on your cloud spend.

No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.

Buyer-side intelligence, monthly.

The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.