So this month I went looking for a link in one of my own posts from early 2013 (some article about, I think, RSS reader alternatives after Google Reader died), and the outbound link was long dead, which, fine, that happens, thats the internet. My move for over a decade has been to just paste the dead URL into the Wayback Machine and grab whatever snapshot exists. Except this time the Wayback Machine wasnt loading right either.
Turns out the Internet Archive got hit with a real one earlier in October. Someone breached them and made off with something like 31 million records worth of emails and hashed passwords, and on top of that a DDoS crew (going by SN_BlackMeta, apparently, which sounds like a Discord server name) has been hammering the site on and off since. archive.org and the Wayback Machine went in and out of read-only mode for stretches, and even now, this close to Halloween, its still not fully back to normal. I read Brewster Kahles updates about it and you can tell theyre running this thing on duct tape and donations, which, I mean, I already knew that, but it hit different when I actually needed it and it wasnt there.
Heres the thing that got me though. I run a blog. Ive been doing it since November 2011, so thats almost thirteen years of posts now, and an embarrassing number of them link out to things that no longer exist in their original form: old Posterous posts (RIP), Google+ threads (RIP), random personal sites that expired and got squatted by ad farms. The Wayback Machine has quietly been the only thing standing between "this old post still kind of makes sense" and "this old post is now forty percent broken links." I never really thought about what happens if that safety net just isnt reliably there anymore. Its one nonprofit, running off donations and a genuinely tiny staff, archiving a meaningful chunk of human memory of the web, and apparently thats also a target now, which is a pretty bleak sentence to type out loud.
I dont think most people who use it daily (students, journalists, randoms like me checking a dead link on a Tuesday night) have any sense of how thin that operation actually is. We treat it like infrastructure, like its just always going to be there the way DNS is always going to be there. It isnt. It's a website with a budget, run by people, and apparently it's also now something a DDoS crew thinks is worth spending effort attacking, for reasons I genuinely dont understand. What is the goal there? Its not like archive.org is sitting on anything ideologically inflammatory, it's a library. Attacking a library is just a weird hill to plant a flag on.
Anyway it made me go do the thing I keep telling myself Ill do and then dont: actually back up techpad properly instead of trusting it'll always be sitting there on whatever host Ive got it on. I ended up moving my nightly database dumps over to a VPS I already had running through Tricknowtech (mostly because git-push deploys there are stupidly easy and I was already using them for domains), and now theres a cron job quietly zipping the whole thing up and shipping it off-site every night instead of me just assuming the Wayback Machine would catch me if things went sideways. Took maybe twenty minutes to set up, which is annoying because it means I really have no excuse for not having done it back in 2019 when I first thought about it.
None of this is really the Internet Archives fault, to be clear. Getting breached sucks, and getting DDoSed on top of it while youre a nonprofit library sucks worse, and I hope the accounts affected werent using reused passwords, though statistically some percentage definitely were. Change your password there if you had an account, is I guess the practical takeaway, and maybe go find one of your own old dead links and see if it still resolves. Mine did, eventually, once things came back up a bit. Not all of them though. Some of the internet from 2013 is just gone now, snapshot or no snapshot, and there's nothing particularly profound to say about that except that it bugs me more than it probably should.