Everybody in my Twitter feed spent this week talking about Steve Ballmer. Fair enough — the guy announced Thursday he's stepping down from Microsoft within the next twelve months, and that's a genuinely big deal for anyone who's spent the last decade making fun of the "developers developers developers" clip. I don't have much to add to that pile. Every tech blog on the internet is going to have a take on Ballmer by Monday, and honestly most of them will say the same three things (stock's been flat for years, tablets caught them flat-footed, he built the enterprise cash cow that's still paying everyone's salary). I'll let them have it.
What actually stopped me cold this week was Thursday's other story, the one that got maybe a tenth of the coverage: Nasdaq just... stopped. For real, in the middle of the trading day.
Around 12:14pm Eastern on August 22nd, quote data for every Nasdaq-listed stock froze up and trading got halted. Not one company, not a sector — the whole exchange. Apple, Microsoft, Facebook, Google, all of it, just sitting there for over three hours while people at the exchange tried to figure out what broke. Markets didn't reopen until about 3:25pm. Reporters immediately started calling it the "Flash Freeze," which is a cute callback to the 2010 Flash Crash but also kind of depressing, since it means we've had enough of these events now that we need a naming convention for them.
The cause, as far as anyone's explained it so far, traces back to the Securities Information Processor, the piece of plumbing that's supposed to merge price quotes coming out of Nasdaq and NYSE Arca into one clean feed that every trading terminal in the world reads from. Something in that handoff choked. Nobody's said publicly yet whether it was a bad software push, a capacity problem, or just bad luck, and I doubt we'll get a real answer for weeks.
Here's the thing that gets me: this is supposed to be the most heavily engineered, most redundant, most stress-tested software on the planet. Billions of dollars move through this stuff every single minute. And it fell over for three hours because of what's basically a message-formatting problem between two systems that are supposed to talk to each other constantly. I was sitting in a coffee shop on 9th Street when the news started trickling out (badly, over spotty wifi that the place still hasn't fixed almost two years after I first complained about it in a post here), and my first reaction wasn't "wow, scary," it was "of course." Of course the plumbing under the entire US stock market is one bad config away from a nap.
I don't even really care about the stock market, to be clear. I don't day-trade, I own basically zero individual stocks, and most of my opinions about finance come from reading Matt Levine-style writeups after the fact. What I care about is that this is the same lesson that shows up over and over in every big system, whether it's an exchange or a website or the power grid: the failure almost never comes from the big flashy piece everyone worries about. It comes from some unglamorous integration layer that three engineers understand and nobody outside the building has ever heard of. Nasdaq has presumably spent millions hardening its actual trading engines against exactly this kind of outage. And the thing that took it down was the boring feed that tells everyone else what the price is.
There's a version of this argument that applies to basically every "modern" piece of infrastructure we've built in the last ten years, cloud included. We keep stacking systems on top of systems, assuming the layer underneath is solid because it's always been solid, until one Thursday it isn't. I don't have a tidy fix for that, and I'm suspicious of anyone who claims they do.
Anyway. Ballmer's leaving, the market forgot how to talk to itself for an afternoon, and somewhere there's a poor engineer at Nasdaq who's going to be in meetings about SIP failover architecture for the rest of the year. Rough week to be that person.