Three Hours Where the Nasdaq Just Stopped

Three Hours Where the Nasdaq Just Stopped

Tech News nasdaq outages tech-infrastructure wall street

So Wall Street just stopped for three hours on Thursday, and I don't think most people outside the finance/tech overlap noticed how weird that actually is.

Around 12:14pm Eastern on August 22nd, trading in every Nasdaq-listed stock just froze. Apple, Google, Microsoft, Facebook, thousands of smaller names, all of it just sat there. Not "prices moved slowly," not "some order types were delayed." Frozen. It stayed that way for something like three hours and eleven minutes before quotes started flowing again around 3:25. For a market that's supposed to execute trades in microseconds, three hours might as well be a geological era.

The cause, once it trickled out, was almost insultingly mundane: a connectivity problem between NYSE Arca and the Securities Information Processor, which is the unglamorous piece of plumbing that's supposed to broadcast a single unified price quote for every stock so nobody's trading on stale numbers. Not a hack. Not a rogue algorithm gone haywire like the Knight Capital mess last year. Just a data feed that stopped talking to another data feed properly, and because the whole system is built to halt rather than let anyone trade on bad information, the entire exchange just... paused.

I keep going back and forth on how to feel about that. On one level it's exactly what you want — the system detected it couldn't guarantee accurate pricing and it stopped rather than let people get burned. That's the responsible failure mode. But the amount of self-congratulation floating around afterward, all this "look how resilient our markets are" talk, rubbed me the wrong way a little. A resilient system that goes down for three hours because one feed hiccuped isn't resilient, it's brittle with a good PR team. Resilient would be the trading continuing on a backup feed without anyone outside the data center noticing. This was closer to duct tape holding until the right person got paged.

And nobody seems to want to say plainly that this is the second time in about thirteen months something like this has knocked a major exchange sideways (Knight Capital's algo disaster cost them $440 million in like 45 minutes back in August 2012, practically the same week on the calendar, which is a coincidence I can't stop thinking about). These are not small companies running this infrastructure. This is supposed to be the most heavily engineered, most heavily tested software on the planet, short of maybe NASA's stuff, and it still falls over.

I'll admit some of my crankiness here is projection. My own site went down for about 30 hours back in March after I tried to migrate the comments database to a new host at 11pm on a Tuesday because "it'll only take twenty minutes." It took twenty minutes and then it took a weekend, because I hadn't tested the migration script against the actual production data volume, just a tiny local copy. Nobody's life savings were riding on techpad staying up, obviously, the stakes are not remotely comparable. But the underlying lesson is the same one every engineer learns the hard way eventually: your system is exactly as reliable as the least-tested path through it, and the least-tested path is usually the one you only hit in production, at the worst possible moment.

(Also, yes, I know everyone's been buried in Ballmer-retiring-from-Microsoft news since yesterday and I'm not going to add another take to that pile today. Enough people are already writing "what does this mean for Microsoft" posts this week. Wall Street forgetting how to trade for three hours felt like the more interesting story nobody was fully sitting with.)

What actually strikes me most is how quietly this got absorbed. Markets reopened the next morning like nothing happened. Trading volume was basically normal by Friday. If a consumer website went dark for three hours mid-afternoon, we'd never hear the end of it — but the infrastructure underneath trillions of dollars in daily transactions gets a day of headlines and then a shrug. Maybe that's fine. Maybe that's just how deeply we've decided to trust systems we don't really understand until they stop working.

I don't have a tidy answer for that one. I just keep thinking about how many other pieces of infrastructure I trust every single day without knowing a single thing about how they actually hold together.