GenAI broke the cloud budget because its cost scales with user behaviour rather than provisioned capacity, and most cost controls were built for capacity. A traditional service costs roughly what you provisioned; a GenAI feature costs more every time a user sends a longer prompt, the model returns a longer answer, or traffic doubles, and that cost lands as per token inference charges, GPU hours, provisioned throughput, and capacity reservations. Because tokens are invisible on a normal dashboard and GPU capacity is bought in large blocks, spend runs ahead of awareness until the monthly bill arrives. The fix is to govern GenAI as its own cost domain: measure cost per token and per outcome, set capacity to a forecast, and attribute spend to the team and feature that drives it.
Here are the drivers, why they evade controls, and the governance that brings them back.
What actually drives GenAI cost?
Four drivers. Token volume, the input and output tokens per request multiplied by request count, which is the core variable of hosted model APIs. GPU hours, the cost of running open weight models on your own accelerators whether or not they are busy. Provisioned throughput, capacity you reserve for guaranteed model performance, billed whether or not you use it. And capacity reservations for scarce GPUs, a commitment that carries the same use it or lose it risk as any other. Each scales differently, and a budget that tracks only one of them misses the rest.
Why does GenAI spend evade normal controls?
Three reasons. First, the unit is invisible: nobody sees tokens the way they see instances, so a prompt change that doubles context doubles cost silently. Second, it scales in real time with demand, so a feature that goes viral or a retry loop that misfires can multiply spend within a day, faster than monthly budgets react. Third, it is bought two ways at once, on demand per token and reserved as GPU capacity or provisioned throughput, so the same workload appears in two very different line items. Capacity bought for a peak that never arrives is pure waste, and on demand token spend left ungoverned compounds quietly.
What governance brings GenAI back in line?
Treat it as a first class cost domain. Instrument token usage per feature and per team so the invisible unit becomes visible. Define a cost per outcome, the spend to serve one resolved ticket, one generated document, one answer, so the business can judge value rather than raw volume. Set GPU capacity and provisioned throughput to a forecast and review utilization, exactly as you would any commitment. Attribute every dollar to a team and feature through tags and labels, and put guardrails on context length, model choice, and retries where cheaper options serve the same outcome.
A worked example
A scaling fintech shipped a GenAI support assistant that was an instant success, and the next cloud bill carried a token charge larger than the team's entire prior compute budget. Instrumenting tokens per conversation showed the assistant sent the full knowledge base as context on every request, and a retry on timeouts was doubling calls. Trimming context to the relevant passages, switching routine queries to a smaller model, and fixing the retry loop cut cost per resolved conversation by more than half, while a right sized provisioned throughput commitment replaced ad hoc on demand peaks. It was part of the program that left the company 41 percent lighter on cloud spend. Figures are verified against billing data and anonymised.
Frequently asked questions
Why is GenAI spend so hard to control?
What are the main GenAI cost drivers?
How do you govern GenAI cost?
Bring GenAI spend back under governance
We help enterprises govern GenAI as its own cost domain, from token instrumentation to capacity strategy, as an independent advisory that takes zero provider commissions. Our guarantee: we reduce your cloud spend or we reimburse our service fee, on a Fixed Fee or a no risk Gainshare basis. Download the AI spend governance guide, read the cloud cost optimization guide, and see AI spend governance for the CFO.
Put a defensible number on your cloud spend.
No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.
The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.