AWS streaming cost on Kinesis Data Streams is driven by four things: the capacity mode you choose, provisioned or on demand, the number of shards in provisioned mode, the volume of data put and retrieved, and how long records are retained. The common trap is provisioning shards for a peak and never scaling them back, so you pay for headroom you stopped using, plus extended retention nobody reads and enhanced fan out consumers nobody needs. Control comes from matching mode to pattern: on demand for spiky or unpredictable traffic so you stop paying for idle shards, provisioned for steady high volume where you can size shards precisely, with retention and consumer choices trimmed to what the workload genuinely requires.
Here is what each lever costs, how to choose capacity mode, and where streaming spend quietly accumulates.
What drives Kinesis cost?
In provisioned mode you pay per shard hour, plus charges for payload units put into the stream, plus extended data retention beyond the default window, plus enhanced fan out if consumers use it. In on demand mode you pay for throughput rather than provisioned shards, and the service scales capacity automatically. Data volume matters, but the structural cost is the capacity you hold and the retention you keep, both of which persist whether or not traffic justifies them. That is why a stream sized for a one off peak stays expensive: the shards and retention do not shrink themselves.
How do you choose between provisioned and on demand mode?
The choice follows the traffic pattern:
- On demand suits spiky, unpredictable, or new workloads. You stop paying for idle shards during quiet periods and avoid the operational work of scaling, at a higher unit rate when fully loaded.
- Provisioned suits steady, high volume streams where throughput is predictable. Sizing shards to real throughput gives a lower unit cost than on demand, provided you actually right size and scale back after peaks.
The expensive mistake is provisioned mode left at a launch peak. If you are not actively managing shard counts against throughput, on demand often costs less in practice because it removes the idle headroom you forget to reclaim.
Where does streaming spend quietly accumulate?
Three places. Retention: every hour of extended retention beyond what consumers replay is pure cost, so trim it to the real recovery window. Enhanced fan out: it gives each consumer dedicated throughput at a per consumer charge, valuable for genuine fan out but wasteful when a shared read would do. Over provisioned shards: capacity sized for a peak that has passed. Review these three on the same cadence as the rest of the bill, because none of them shrinks on its own and all of them compound month over month.
A worked example
A scaling fintech provisioned Kinesis shards generously ahead of a product launch and never revisited them, while extended retention sat at a multi day window that no consumer replayed beyond a few hours. Moving the spikiest streams to on demand mode, right sizing shards on the steady streams to measured throughput, and cutting retention to the real replay window reduced streaming cost sharply with no loss of records or throughput. Reclaiming this quiet line item was one of many moves in the program that left the company 41 percent lighter on cloud spend. Figures are verified against billing data and anonymised.
Frequently asked questions
What makes Kinesis expensive?
Should I use on demand or provisioned Kinesis?
How does retention affect streaming cost?
Bring your streaming spend under control
We help AWS teams match streaming capacity mode to traffic, right size shards and retention, and reclaim the quiet line items, across AWS, Azure, GCP, and OCI. We take zero provider commissions and answer only to you, on a Fixed Fee or a no risk Gainshare basis, with a guarantee that holds us to it: we reduce your cloud spend or we reimburse our service fee. Book a strategy call, and follow more in The Cloud Spend Navigator.
Put a defensible number on your cloud spend.
No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.
The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.