Last Tuesday morning I was trying to post something dumb on X and it just wouldn't load. Fine, happens. Then I went to check if it was just me and Downdetector wouldn't load either. That's the part I keep coming back to a week later, not the outage itself but the fact that the site whose entire job is telling you "yes, X is down" was also down, because it turns out even Downdetector leans on Cloudflare somewhere in its stack. There's something almost too neat about that. Like the fire department's phone line running through the building that's on fire.
For anyone who missed it: Cloudflare had a rough morning on November 18th, starting somewhere around 6:20am Eastern, and it rippled out to a genuinely absurd list of sites. X was flaky, ChatGPT was throwing errors, Spotify wouldn't play, League of Legends players got kicked mid-match, Canva stalled out. I don't even use half of those regularly and I still noticed, because my group chat turned into a live incident report for about two hours. Somebody's theory was a DDoS attack, which is always the first guess and is almost never right anymore. This one traced back to a config file for their bot-management system that grew past a size limit and started choking their systems when it got pushed out. A permissions change on a database query, apparently, changed what data got pulled into that file. Not a hack. Not even close to a hack. Just a file that got too big.
I only really cared because I was mid-scroll, and this blog itself didn't blink, because techpad runs off a cheap little VPS with no CDN in front of it at all, which normally I think of as a minor embarrassment (no image optimization, no edge caching, load times that are fine but not impressive) and for once felt like a small victory. Nothing to route through. Nothing to fail.
What actually gets me is that this is the second time in a month. AWS had its own mess on October 20th, a DNS resolution problem tangled up with DynamoDB in their us-east-1 region, and that one hit an even weirder spread of stuff: Snapchat, Fortnite, Venmo, Alexa, Ring doorbells. My father-in-law couldn't get his smart lock to open that morning and just stood on his own porch for ten minutes waiting for it to sort itself out, which I still think about, because at some point we agreed it was fine for the front door of a house to depend on a data center in Virginia. Nobody voted on that. It just happened, one convenient integration at a time, and now two outages in five weeks have shown pretty plainly that a shocking amount of the internet, and apparently some amount of the physical world, sits on maybe three companies' infrastructure.
I'm not going to pretend I have a fix. Multi-cloud redundancy is expensive and most companies aren't going to spend real money hedging against a few hours of downtime a year, the math doesn't work for them even if it's annoying for the rest of us. And honestly the alternative isn't obviously better. Before everyone centralized onto Cloudflare and AWS and a handful of others, sites just fell over on their own all the time, individually, for dumber reasons, with worse security. At least this way when it breaks it breaks all at once and you find out fast.
What bugs me is the framing every time this happens, this "internet is down" language from big outlets like it's some cosmic event, when what actually happened is three or four specific companies had a bad morning and a huge number of unrelated products turned out to be quietly dependent on them. That's not the internet being down. That's a concentration problem, and it keeps getting reported like weather.
Anyway. X came back, Downdetector came back to tell everyone X was fine now, and by lunchtime the group chat had moved on to arguing about something else entirely. I did not move my blog to Cloudflare that day, and I'm not going to pretend that was some principled infrastructure decision instead of just laziness that happened to pay off once.