TL
The short answer

Disaster recovery on a budget is a tiering exercise: set recovery time and recovery point objectives per workload, then pick the cheapest of four patterns that meets them, which are backup and restore, pilot light, warm standby, and active active. The cost driver is how much idle capacity you keep running, and it rises steeply as recovery time approaches zero, so a single tier applied to everything always overpays. Cross cloud DR adds egress charges and a second platform to operate, and it is justified only when the risk being insured is provider level or concentration risk, not regional failure that a second region on the same provider covers more cheaply.

The discipline is to insure each workload against the right risk at the right price. Here are the patterns, the cost mechanics across AWS, Azure, GCP, and OCI, and how to decide.

What are the DR patterns and what do they cost to run?

There are four standard patterns, and they differ almost entirely in how much you pay to keep capacity ready before a failure ever happens. The data is always replicated; what changes is the compute and the readiness.

PatternTypical recovery timeStanding cost driver
Backup and restoreHoursStorage and egress only; compute rebuilt on demand from infrastructure as code
Pilot lightTens of minutesStorage, plus a minimal core (databases, key services) kept running small
Warm standbyMinutesA scaled down but always on copy of the full stack
Active activeNear zeroFull duplicate capacity serving or ready to serve traffic

Cost climbs roughly in that order, and the jump from warm standby to active active is the steepest because you are paying for a second full estate. The buyer mistake is defaulting the whole portfolio to warm standby or active active for operational comfort. Most workloads do not need it.

How do RTO and RPO set the budget?

Recovery time objective is how long you can be down. Recovery point objective is how much data you can afford to lose. Tighter objectives force warmer, more expensive patterns. The single most effective cost move in DR is to set these per workload from the business impact, not to inherit one tier for everything. A regulated transaction system may genuinely need near zero recovery, but the internal reporting database almost certainly tolerates a few hours, and backup and restore is a fraction of the cost.

Write the objectives down with the workload owner and the business owner together, because engineering tends to gold plate recovery and finance tends to underfund it. The agreed number is what you size to, and revisiting it annually catches workloads whose importance has changed.

The buyer test

For each workload, ask what a four hour outage actually costs the business. If the honest answer is modest, you are overpaying for any pattern warmer than backup and restore. Pay for recovery speed only where downtime is genuinely expensive.

Where do the cross cloud costs hide?

Running DR on a different provider than production sounds like the strongest insurance, but it carries real costs that same cloud DR does not. Data egress between providers is charged on the way out, and DR means continuously replicating data across that boundary, so egress becomes a standing monthly line rather than a one time move. You also operate two platforms, which means two sets of skills, two tooling stacks, and two security models to keep current. Per cloud, egress economics differ: OCI prices egress materially cheaper than the hyperscalers, AWS NAT gateways and cross region transfer quietly add up, Azure and GCP each meter cross region and internet egress on their own schedules. Those differences can swing a cross cloud DR design by a wide margin.

The honest framing is that cross cloud DR insures against provider level failure and concentration risk. A second region on the same provider, which avoids inter provider egress, insures against regional failure for much less. Choose cross cloud only when the risk you are actually worried about is the provider itself, a contractual concentration limit, or a regulatory requirement to avoid single vendor dependence.

How do you bring DR spend down without losing resilience?

Start by retiering. Map every workload to objectives and move anything overspecified down a pattern. Then attack the standing costs of whatever pattern remains. Use cheaper storage classes and lifecycle policies for backups and replicas, since DR data is rarely read. Keep pilot light and warm standby cores genuinely small and scale them out only on failover. Reserve commitments only for capacity that truly runs all the time, never for standby you would scale from zero, because committing to idle DR capacity is paying twice. Finally, test failover on a schedule. Untested DR is a cost with no proven benefit, and tests routinely reveal standby capacity that was larger than the recovery actually needs.

Worked example

A Fortune 500 retailer ran warm standby for its entire estate across two clouds for operational simplicity. Setting objectives per workload showed that most internal systems tolerated multi hour recovery. We moved those to backup and restore with lifecycle managed replicas, kept warm standby only for the customer facing transaction path, and replaced a cross cloud replica with a same provider second region where the risk was regional. The resilience the business actually required was preserved, and the DR estate was a meaningful part of the work that left total cloud spend materially lighter. Figures are verified against billing data and anonymized.

Where this fits in your multicloud program

DR design is one of the places multicloud architecture and cost meet head on, alongside networking and lock in decisions. See how inter provider data movement is priced in multicloud networking costs explained, and how to weigh resilience against vendor dependence in avoiding lock in without overpaying. The full cross cloud picture lives in the cross cloud cost optimization guide.

Frequently asked questions

What is the cheapest DR pattern across clouds?
Backup and restore. You replicate data to a second region or cloud and rebuild on demand from infrastructure as code, paying for storage and egress rather than idle compute. It fits workloads that tolerate a recovery time of hours.
Does cross cloud DR cost more than same cloud DR?
Usually, because inter provider egress is charged and you operate two platforms. It earns the premium only when the risk is provider level or concentration risk, not regional failure, which a second region on the same provider covers more cheaply.
How do RTO and RPO drive DR cost?
Tighter objectives require warmer standby and more idle capacity, and cost rises steeply as recovery time approaches zero. Setting objectives per workload rather than one tier for everything is the biggest DR cost lever.

Right size your cross cloud resilience

We help enterprises tier disaster recovery to real recovery objectives and strip out standby they pay for but never use, across AWS, Azure, GCP, and OCI. Our guarantee: we reduce your cloud spend or we reimburse our service fee. Pricing is either a Fixed Fee scoped up front or Gainshare, a share of verified savings with no retainer and no risk.

Independent · buyer-side

Put a defensible number on your cloud spend.

No provider in the room, no published price list. Tell us your footprint and we will scope the savings against your billing data — we reduce your cloud spend or we reimburse our service fee.

Buyer-side intelligence, monthly.

The Cloud Spend Navigator: what changed in cloud pricing, commitments, and FinOps — no vendor spin.