TL
The short answer

Cost per AI outcome is the infrastructure cost of producing one unit of business value, such as a resolved support ticket, a generated document, or a completed recommendation. You compute it by attributing the underlying token, GPU, provisioned throughput, and capacity reservation costs to a feature, then dividing by the count of outcomes that feature produced. The metric matters because total AI spend is unbounded and unrevealing, while cost per outcome tells you whether a feature is getting cheaper or more expensive as it scales, which model and architecture choices are paying off, and where governance should focus. Measuring it well depends on clean allocation of AI spend down to the feature, which most estates lack by default.

Here is the metric defined, the inputs it needs, and how to act on it once you have it.

Why measure cost per outcome instead of total AI spend?

Total AI spend rises with adoption whether or not the spend is efficient, so it cannot tell you if you have a problem. Cost per outcome can. If the cost to resolve one ticket with an assistant is falling as volume grows, scale is working in your favour. If it is flat or rising, you are buying growth at a cost the business may not have agreed to. The metric also makes model and architecture choices legible: a cheaper model that needs more retries, or a retrieval design that doubles token use, shows up immediately in cost per outcome where it hides in the aggregate.

What goes into the calculation?

The numerator is the fully attributed infrastructure cost of the feature. For a typical generative feature that includes:

  • Token cost. Input and output tokens at the model's rate, including the context you send and any retries.
  • Serving cost. Provisioned throughput or on demand inference, or the GPU hours behind a self hosted model, plus any capacity reservation you hold for headroom.
  • Supporting infrastructure. Vector database queries, retrieval, caching, and the orchestration compute around the model call.

The denominator is the count of outcomes, defined as the unit the business funds: a resolved conversation, a generated summary, a scored lead. The discipline that makes this possible is the same allocation discipline the rest of the cloud bill needs, tags and billing exports that carry the feature identity, applied to AI line items across AWS, Azure, GCP, and OCI.

How do you act on the number once you have it?

  • Set a target and watch the trend. A cost per outcome that falls with volume is the signal that scale is healthy. A rising one is a governance trigger.
  • Compare model and prompt choices on equal terms. The cheapest token rate is not always the cheapest outcome once retries and context are counted.
  • Match serving to demand. Provisioned throughput pays off under steady load; on demand inference wins for spiky or early stage features. The crossover is visible in cost per outcome.
  • Fund features on unit economics, not enthusiasm. A feature with a clear, improving cost per outcome is easy to fund; one without a measured number is a blank cheque.

A worked example

Worked example

A scaling fintech ran a support assistant whose monthly bill kept climbing, and leadership could not tell whether that was success or waste. We attributed token, inference, and retrieval cost to the feature and divided by resolved conversations to get a cost per resolution. The number exposed that a large share of spend came from oversized context sent on every call and from retries on a model that was cheaper per token but less reliable. Trimming context and switching to a model with a better outcome cost, not a better token rate, lowered the cost per resolution while volume kept growing. The unit metric, not the total, made the fix obvious and fundable. Figures are verified against billing data and anonymised.

Frequently asked questions

What is cost per AI outcome?
It is the cloud infrastructure cost of producing one unit of business value from an AI feature, such as a resolved ticket or a generated document. You attribute token, serving, and supporting costs to the feature and divide by the number of outcomes it produced.
Why is total AI spend a poor metric?
Because it rises with adoption whether or not the spend is efficient, so it cannot tell you if a feature is getting cheaper or more expensive as it scales. Cost per outcome reveals that trend and makes model and architecture choices comparable.
What do you need to measure it accurately?
Clean allocation of AI spend down to the feature: tags and billing exports that carry feature identity across token, inference, GPU, and supporting line items on AWS, Azure, GCP, and OCI, plus a clear definition of the outcome the business funds.

Put a unit cost on every AI feature you run

We help enterprises define and measure cost per AI outcome across AWS, Azure, GCP, and OCI, so token, GPU, and inference spend maps to value the business can govern. Independent and buyer side, with zero provider commissions. Our guarantee: we reduce your cloud spend or we reimburse our service fee, on a Fixed Fee or no risk Gainshare basis. Book a strategy call, and follow more in The Cloud Spend Navigator.

Independent · buyer-side

Put a defensible number on your cloud spend.

No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.

Buyer-side intelligence, monthly.

The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.