To forecast AI spend under growth, build the forecast from the drivers that actually move it: request and token volume, cost per call by model and feature, and the GPU or provisioned throughput hours behind them. Project each driver from a business volume the company already forecasts, such as active users or documents processed, and convert it to spend through a unit cost. Then split the result into a stable floor that runs every day and a variable tail that spikes with launches, cover the floor with reservations, provisioned throughput, or committed use, and let the tail ride on demand or spot. The instruments differ by cloud, with capacity reservations and provisioned throughput on the hyperscalers and Universal Credits on OCI, but the discipline is the same: commit to a defensible floor, never to a peak you hope to reach.
AI is the fastest growing line on most 2026 cloud bills, and it behaves unlike steady infrastructure. Volumes can double in a quarter, a model swap can halve or triple cost per call, and capacity comes in coarse blocks. A forecast that ignores those dynamics fails the moment growth arrives. Here is how to build one that holds.
Why does AI spend resist normal forecasting?
Traditional cloud forecasting leans on a stable base that grows gently, so a trend line plus a margin works. AI breaks all three assumptions. Usage is adoption driven and rises nonlinearly as features land and habits form. Cost per call is not fixed, because model choice, prompt size, context length, and caching all move it, and teams change those constantly. And capacity is lumpy: GPU reservations and provisioned throughput come in blocks you either fill or waste. The aggregate bill is the product of these moving parts, so forecasting the aggregate directly is forecasting noise. You have to model the parts.
What drivers should the forecast model?
Three layers. The demand layer is the business volume that generates calls, such as active users, sessions, or documents, which the company usually already projects. The intensity layer converts demand into AI work: calls per user, tokens per call, and how often a request hits an expensive path such as a large context window or a premium model. The cost layer turns work into money through a unit cost per thousand tokens, per request, or per GPU hour. Forecasting each layer separately means a change in any one, a new feature, a model swap, a caching win, shows up as a specific, defensible adjustment rather than a mystery in the total.
What unit metric makes the forecast legible?
Pick a unit that ties spend to value: cost per request, per active user, per processed document, or per thousand tokens. A unit metric does two jobs. It lets finance forecast AI spend straight from a business number they already plan, so the forecast moves with the plan automatically. And it reveals direction of travel, because a unit cost that falls as you scale means efficiency is winning, while one that rises means growth is outrunning optimization and needs attention before the absolute number alarms anyone. The unit metric is what turns an AI forecast from a finance argument into an engineering signal.
| Forecast layer | What you model | Source |
|---|---|---|
| Demand | Active users, sessions, documents | Existing business plan |
| Intensity | Calls per user, tokens per call, premium path rate | Usage telemetry |
| Cost | Cost per thousand tokens, per request, per GPU hour | Provider pricing, verified |
| Capacity | Stable floor versus variable tail | Utilization history |
Table: the four layers of a driver based AI spend forecast and where each input comes from.
How do you commit when the future is uncertain?
Commit to the floor. Even fast growing AI workloads have a baseline that runs every day, plus a variable tail that spikes with launches, experiments, and seasonal demand. Size reservations, provisioned throughput, or committed use to the floor you are confident will persist, and let the uncertain tail ride on demand or, for fault tolerant batch and training, on spot and preemptible capacity. This captures the discount on the steady part without locking you into capacity that growth might route elsewhere or that a model change might make obsolete. The forecast feeds the commitment: a credible floor, not an optimistic peak, sets the coverage. Treat any provider list price you build from as indicative and verify it against the current pricing page before committing.
A worked example
A European SaaS company added an AI assistant and watched inference spend triple in a quarter, with finance unable to forecast the next one. Rebuilding the forecast from drivers showed that most growth came from tokens per call rising as prompts grew, not from more users, so the team trimmed context and added caching, and cost per request fell even as volume climbed. The forecast then separated the steady serving baseline from launch driven spikes, covered the baseline with provisioned throughput and committed capacity, and left the spikes on demand. Spend became predictable a quarter ahead, and the optimization work made the AI line materially lighter against the original trajectory. Figures are verified against billing data and anonymised.
Where this fits the wider review
A forecast is only as good as the cadence that maintains it, covered in the AI cost review cadence, and it depends on the telemetry described in observability for AI cost. To present the result upward, use the framing in AI unit economics the board understands. The full cross cloud picture, including how AI sits alongside every other lever, lives in the cross cloud cost optimization guide.
Frequently asked questions
Why is AI spend so hard to forecast?
How do you commit to AI capacity when usage is uncertain?
What unit metric should an AI forecast use?
Build your AI forecast with us
We build driver based AI spend forecasts and size the commitments behind them across AWS, Azure, GCP, and OCI, with zero provider commissions. Our guarantee: we reduce your cloud spend or we reimburse our service fee. Pricing is either a Fixed Fee scoped up front or Gainshare, a share of verified savings with no retainer and no risk.
Put a defensible number on your cloud spend.
No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.
The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.