On demand inference charges per token consumed and provisioned throughput charges a fixed rate for reserved model capacity, so on demand is cheaper for variable or low volume traffic and provisioned throughput is cheaper once sustained demand keeps the reserved capacity highly utilized. Provisioned throughput also guarantees latency and capacity that on demand endpoints can throttle under load, so high volume production traffic often chooses it for predictability as well as price.
Here is how each model bills, where the crossover sits, what provisioned throughput buys beyond price, and a worked example of the decision for a real traffic pattern.
How does each pricing model bill?
On demand inference, the default on every managed model platform, charges per thousand input and output tokens with no commitment. You pay only for what you send, and the rate per token is the same whether you send one request a day or thousands a second, until you hit a throughput quota and requests start to throttle. Provisioned throughput instead reserves a fixed block of model capacity, billed by the hour for a monthly or longer term, and within that block you can run as many tokens as the capacity supports at no extra per token charge. AWS Bedrock sells this as Provisioned Throughput, Azure OpenAI as Provisioned Throughput Units, and Google Vertex AI as provisioned throughput, each with monthly or annual commitments that lower the effective hourly rate.
Where is the crossover?
The crossover is the volume at which the fixed reserved rate equals what you would pay per token on demand for the same work. Below it, on demand is cheaper because you are not paying for idle reserved capacity; above it, provisioned throughput is cheaper because the reserved block absorbs more tokens at no marginal cost. The reserved capacity only pays off if you keep it busy, so the real question is your sustained utilization across the day, not your peak. Traffic that runs hot for two hours and quiet for twenty two rarely justifies a full day of reserved capacity; traffic that holds a high floor around the clock crosses over quickly.
| Traffic pattern | Cheaper model | Why |
|---|---|---|
| Low or unpredictable volume | On demand | No idle reserved capacity to pay for |
| Spiky, with long quiet periods | On demand, or a small reserved floor | Reserved capacity sits idle between spikes |
| High and steady around the clock | Provisioned throughput | Reserved block stays busy, no per token charge above it |
| High floor plus spikes | Reserved floor plus on demand burst | Commit the floor, burst the rest per token |
What does provisioned throughput buy beyond price?
Cost is only half the decision. Provisioned throughput reserves capacity, so it delivers predictable latency and a guaranteed request rate that on demand endpoints cannot promise once a region is busy. For interactive products where a slow or throttled response costs conversions, that guarantee can justify reserved capacity even slightly before the pure price crossover. The mirror risk is utilization: reserved capacity you do not use is pure waste, the same use it or lose it shape as a compute commitment, so size the reservation to a defensible forecast of sustained demand rather than to the peak you hope to reach.
A worked example
A software company ran a customer facing assistant entirely on demand. Daytime traffic held a high, steady floor for roughly ten hours, then fell away overnight. Modelling the floor showed that reserving capacity to cover the steady daytime band, while leaving the overnight tail and occasional spikes on demand, cost less than paying per token for the whole floor, and removed the throttling that had been hurting response times at peak. Committing the full twenty four hours would have stranded capacity overnight and cost more. The decision was to reserve the predictable band and burst the rest. Figures are verified against billing data and anonymised; rates are indicative and depend on model, region, and term.
Frequently asked questions
Is provisioned throughput cheaper than on demand inference?
When should I reserve inference capacity?
What is the risk of provisioned throughput?
Right size your inference commitment
We model on demand against provisioned throughput across Bedrock, Azure OpenAI, and Vertex, find your crossover, and size reservations to a forecast you can defend, as an independent advisory that takes zero provider commissions and answers only to you. Our guarantee: we reduce your cloud spend or we reimburse our service fee, on a Fixed Fee or a no risk Gainshare basis. Download the cloud cost optimization playbook, read the cross cloud cost optimization guide, and on the unit math see token economics for enterprise buyers. For monthly buyer side analysis, subscribe to The Cloud Spend Navigator.
Put a defensible number on your cloud spend.
No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.
The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.