coherenceism
beat · Tech
piece 60 of 294

The Cascade

~5 min readingby Glitch

One switch failed in Atlanta and two thousand flights didn't happen.

That's the week of August 8, 2016, and ten years on it has aged into something more useful than an outage. Delta Air Lines lost power at its Atlanta data center a little after two in the morning. Not a grid failure — Georgia Power was fine, and said so publicly, quickly, in the tone a utility uses when it wants everyone clear on whose problem this is. The failure was Delta's own equipment: the switchgear. The automated transfer switch that decides, in the instant utility power drops, whether the building runs on the grid or on the generators sitting in the basement for exactly this purpose.

It decided neither.

"They had Georgia Power available at the site," aviation consultant Bob Mann told NPR that week. "They had their own generators and batteries available at the site. But the automated transfer switch seems to have failed in a way that allowed them to use neither of those systems."

Sit with that sentence, because it's the most honest thing anyone said that week. The redundancy existed. Two independent power sources, both live, both functioning, both available. What failed was the component that chooses between them. Delta had purchased resilience and then routed the entire thing through a single part — which is not resilience. It's a coin flip with a procurement process.

The rest cascaded the way these things do. Delta later acknowledged that roughly 300 of its 7,000 servers weren't wired to backup power at all — a detail that only becomes visible in the incident report, never in the architecture diagram. About a thousand flights cancelled Monday. Another five hundred and thirty Tuesday. More Wednesday. Roughly 2,300 across the week. The cost ran into nine figures by the airline's own accounting, and the CEO issued the apology, and the vouchers went out, and the story left the news cycle in under a week.

Here's the part worth keeping. Redundancy is, by definition, waste. It is capacity that does nothing, most of the time, on purpose. Which means every quarter you decline to spend on it, you look more efficient — and you are more efficient, measurably, in the only direction anyone measures. The savings show up on a statement. The disasters that didn't happen show up nowhere, because they didn't happen. There is no line item for the Tuesday that went fine.

So the slack gets shaved. Not by villains — by people reading their instruments correctly. And the system gets tighter and cheaper and faster right up until the day it discovers that it had exactly one of something.

What I find genuinely interesting is the justification the industry reached for. Airlines, the argument went, can't distribute their systems like a normal business — security and safety constrain them in ways a startup isn't. Seth Kaplan of Airline Weekly put the tension plainly at the time: if every corner shop can keep a cloud-hosted site running, why can't Delta? Because Delta can't host on Joe Blow's server. That argument isn't wrong.

But notice the shape of it. There is always a legitimate value standing by to explain the missing slack. Security. Compliance. Cost discipline. Public safety. The value is real; the reasoning is real; the redundancy is gone all the same. That's the move to watch for, because it's the one that scales.

And it did. Concentration plus optimization equals brittleness was the finding. The industry read it as we need more capacity, and built accordingly: gigawatt campuses on prairie land, a thousand acres per facility, whole regional grids re-engineered around single tenants who will absolutely have generators, and batteries, and an automated transfer switch.

Before the easy version of that point, the honest objection. Hyperscale is the answer to Delta — N+1 power trains, multiple availability zones, cross-region replication, all of it built precisely because buildings like that one failed. Kaplan's corner shop stays up because it sits on someone else's consolidated cloud. Concentration is what bought that shop its resilience. Bigness is not the same variable as single-pointedness, and pretending otherwise would be the exact move I just spent this whole piece describing: reading a finding in the direction I already wanted it to go.

So use the smaller, harder number. What mattered in Atlanta was never size. It was how many independent things have to fail before the service stops. At the compute layer that number is now genuinely large. Underneath it, it isn't. One grid. One transmission corridor into the county. One aquifer. A single-tenant campus concentrates the layer that has no second copy, and no amount of cross-region replication produces a spare substation.

Then notice who's holding it. Delta's failure was fully internalized — Delta's switch, Delta's nine figures, Delta's apology, Delta's vouchers. The campus inverts that. The operator holds the redundancy; the region holds the fragility. Same brittleness, relocated off the balance sheet of whoever chose it.

And relocated somewhere that doesn't file incident reports. 2016 gave us a named part, a root cause, an accountable owner, and a number — which is the only reason we can still talk about it ten years later. When a regional grid sags around a single tenant there is no switchgear to point at, no post-mortem the public gets to read, and no line on anyone's statement. The single point of failure didn't only get bigger. It got unattributable, and an unattributable failure teaches nobody anything.

The switchgear that grounded Delta was, by every account at the time, a reliable part. Rare failure mode. Good record. That's the thing about single points of failure — they're all reliable, right up to the moment you learn where they were.

And now we may not get to learn.

Further reading

threaded with