The Ratio

Our weekly newsletter on reliability economics.

I run a benchmark that nobody asked for. 121 enterprise teams have taken it anyway. Every Tuesday I send you the one number that surprised me and the seven links that explain why it matters to me.

The newsletter is how I think out loud about what the data says.

No sponsors. No AI slop. Hit reply any time — I read everything.
Prefer RSS? reliabilityeconomics.com/blog/feed/the-ratio.xml

Issue №21October 6, 2026

The Ratio

A weekly newsletter on reliability economics


The Number

27 of 58 tech companies

Nearly half of all Technology organizations in the benchmark are classified as over-investing in reliability prevention — almost as many as the 30 classified as under-investing.

58 tech companies in the benchmark. 27 are over-investing in reliability. 30 are under-investing. 1 is near-optimal. One.

The Ratio

Tech doesn't have a spending problem. It has an aiming problem. The odds of overshooting are almost identical to the odds of undershooting, which is what you'd expect from an industry that treats reliability budgets like a fixed line item instead of a function of what failures actually cost. This is the equivalent of an insurance company pricing every policy the same regardless of claims history. The premium feels rational until the loss hits. Across 58 organizations in the benchmark, exactly one landed in the right range. That's not variance. That's an industry guessing.

Nearly half of tech companies over-invest in reliability. The problem isn't too little spending. It's that almost nobody spends right.



The Crowd Favorite

  1. Paranoid Android — Radiohead ↗ — A service that changes behavior mid-request without logging the state transition turns every post-incident review into guesswork. You're not investigating. You're reconstructing.

  2. Harder, Better, Faster, Stronger — Daft Punk ↗ — Scale throughput without scaling your error budgets and the failure rate grows with the traffic. Faster and broken is still broken.

  3. Message In A Bottle — The Police ↗ — An alert without acknowledgment confirmation means the on-call engineer finds out about the page at escalation, not when the fire starts.

  4. Learning to Fly — Pink Floyd ↗ — Pushing to production without automated canary health gates is flying without instruments. The ground shows up before the altimeter does.

  5. Where The Streets Have No Name - Remastered — U2 ↗ — A deployment with no rollback path turns every bad push into a forward-redeploy incident. Whoever is on-call owns it now.

Prevention

Five gaps before they page someone


The Challenger — Book Review

Release It! Second Edition — Michael T. Nygard (Pragmatic Programmers)

Maps stability anti-patterns, cascade failures, unbounded queues, blocked threads, to specific blast radii. The takeaway for anyone who cares about cost: anti-patterns aren't bugs. They're deferred cost. Each one moves money from the prevention budget to the incident response budget. The circuit breaker chapter alone pays for the time.

Read if: you own a distributed system or an on-call rotation and can't name the pattern behind your last three outages.

Skip if: every integration point already carries a timeout, a bulkhead, and a documented failure class.

Prevention

Anti-patterns are deferred incident cost


The Ratio is a weekly newsletter by Florian Hoeppner.

Take the assessment → reliabilityeconomics.com/benchmark
Reply to this email with your take.

The Ratio

Our weekly newsletter on reliability economics.

I run a benchmark that nobody asked for. 121 enterprise teams have taken it anyway. Every Tuesday I send you the one number that surprised me and the seven links that explain why it matters to me.

No sponsors. No AI slop. Hit reply any time — I read everything.
Prefer RSS? reliabilityeconomics.com/blog/feed/the-ratio.xml