The Ratio
A weekly newsletter on reliability economics
The Number
27 of 58 tech companies
Nearly half of all Technology organizations in the benchmark are classified as over-investing in reliability prevention — almost as many as the 30 classified as under-investing.
58 tech companies in the benchmark. 27 are over-investing in reliability. 30 are under-investing. 1 is near-optimal. One.
Tech doesn't have a spending problem. It has an aiming problem. The odds of overshooting are almost identical to the odds of undershooting, which is what you'd expect from an industry that treats reliability budgets like a fixed line item instead of a function of what failures actually cost. This is the equivalent of an insurance company pricing every policy the same regardless of claims history. The premium feels rational until the loss hits. Across 58 organizations in the benchmark, exactly one landed in the right range. That's not variance. That's an industry guessing.
Nearly half of tech companies over-invest in reliability. The problem isn't too little spending. It's that almost nobody spends right.
This Week in Reliability
AI Velocity Meets Human Burnout
The productivity gains from AI-assisted development are creating a second-order reliability crisis: faster shipping, more incidents, and SRE teams pushed past sustainable limits. This week surfaces the hidden cost of acceleration—burnout isn't a people problem, it's a systems economics problem.
Deep Reads
Burning the Candle at Both Ends: The Human Cost of the AI Velocity Paradox
Harness Blog · Primary evidence
Research shows AI is enabling engineers to ship faster than ever, but the same acceleration is pushing teams to the breaking point. The two trends are described as inseparable.
This is the first real data point naming what practitioners already feel: AI closes the velocity gap but opens a sustainability chasm. Faster deploys don't reduce toil—they multiply the surface area for failure and the cognitive load on the humans who catch it.
Speed without guardrails moves the burnout curve left.
Quick thoughts on GitHub's Sept 23 incident
Lorin Hochstein · Symptom of acceleration fatigue
GitHub experienced elevated 500 and 404 responses on Sept 23. The incident write-up is very short, and Ole Peder Brandtzæg flagged it for review.
When a platform this critical ships a two-paragraph postmortem, it signals either incident fatigue or a normalization of complexity outpacing investigative capacity. Either way, it's a canary: the pace of change is outrunning our ability to learn from failure.
Incidents too frequent to warrant deep investigation.
What managing emergencies at Burning Man teaches me about incident management in tech
Brent Chapman · Adjacent metaphor—high-tempo coordination
The Burning Man Emergency Services Department runs a CAD system and municipal-grade radio dispatch around the clock in conditions of spotty cell coverage and high complexity.
When your incident response looks more like disaster relief than software operations, you've crossed into unmeasured operational debt—this is what sustained high-tempo environments do to coordination systems.
Incident management becomes disaster response at scale.
Fixing Nginx 502 Bad Gateway
Francis Morkeh Mensah · Toil multiplier from velocity
A tiny Nginx proxy buffer misconfiguration caused a production login outage. The fix was four lines.
The gap between root cause simplicity and discovery complexity is where engineer hours vanish—this is the toil tax of velocity without instrumentation depth.
Four-line fix, hours of firefighting.
First-Class Citizens: InMobi's Senior Technical Lead on Fixing Stale Runbooks
StackGen · Prevention gap from acceleration
InMobi's senior technical lead discusses treating runbooks as first-class citizens to address the problem of stale operational documentation.
Runbook drift is burnout written down—when your documentation can't keep pace with your deploy velocity, every incident becomes a knowledge archaeology expedition.
Documentation velocity lags code velocity.
How SRE Architects Balance Reliability, Scalability, and Cost
SRE School · Traditional prevention under pressure
SRE architects automate toil away as a rule—if a task happens more than twice, they script it. This prevents human mistakes and keeps teams focused on strategic goals.
The 'automate after twice' heuristic only works when you have time to write the script before incident three—AI velocity breaks that assumption.
Automation rules assume time you don't have.
The Crowd Favorite
-
Paranoid Android — Radiohead ↗ — A service that changes behavior mid-request without logging the state transition turns every post-incident review into guesswork. You're not investigating. You're reconstructing.
-
Harder, Better, Faster, Stronger — Daft Punk ↗ — Scale throughput without scaling your error budgets and the failure rate grows with the traffic. Faster and broken is still broken.
-
Message In A Bottle — The Police ↗ — An alert without acknowledgment confirmation means the on-call engineer finds out about the page at escalation, not when the fire starts.
-
Learning to Fly — Pink Floyd ↗ — Pushing to production without automated canary health gates is flying without instruments. The ground shows up before the altimeter does.
-
Where The Streets Have No Name - Remastered — U2 ↗ — A deployment with no rollback path turns every bad push into a forward-redeploy incident. Whoever is on-call owns it now.
Five gaps before they page someone
The Challenger — Book Review
Release It! Second Edition — Michael T. Nygard (Pragmatic Programmers)
Maps stability anti-patterns, cascade failures, unbounded queues, blocked threads, to specific blast radii. The takeaway for anyone who cares about cost: anti-patterns aren't bugs. They're deferred cost. Each one moves money from the prevention budget to the incident response budget. The circuit breaker chapter alone pays for the time.
Read if: you own a distributed system or an on-call rotation and can't name the pattern behind your last three outages.
Skip if: every integration point already carries a timeout, a bulkhead, and a documented failure class.
Anti-patterns are deferred incident cost
The Ratio is a weekly newsletter by Florian Hoeppner.
Take the assessment → reliabilityeconomics.com/benchmark
Reply to this email with your take.