Cloudfare outage post mortem

homura1650@lemmy.world · 18 hours ago

Cloudfare outage post mortem

floquant@lemmy.dbzer0.com · 14 hours ago

Their motivation is that that file has to change rapidly to respond to threats. If a new botnet pops up and starts generating a lot of malicious traffic, they can’t just let it run for a week

unexposedhazard@discuss.tchncs.de · edit-2 13 hours ago

How about an hour? 10 minutes? Would have prevented this. I very much doubt that their service is so unstable and flimsy that they need to respond to stuff on such short notice. It would be worthless to their customers if that were true.

Restarting and running some automated tests on a server should not take more than 5 minutes.

SMillerNL@lemmy.world · 11 hours ago

5 minutes of uninterrupted DDoS traffic from a bot farm would be pretty bad.

ramble81@lemmy.zip · edit-2 9 hours ago

5 hours of unintended downtime from an update is even worse.

Edited for those who didn’t get the original point.

SMillerNL@lemmy.world · 9 hours ago

It wasn’t an unintentional update though, it was an intentional update with a bug.

ramble81@lemmy.zip · 9 hours ago

Edited. My point still stands.

dafta@lemmy.blahaj.zone · 10 hours ago

Significantly better than several hours od most of the internet being down.

SMillerNL@lemmy.world · 9 hours ago

Maybe not updating bot mitigation fast enough would cause an even bigger outage. We don’t know from the outside.

Echo Dot@feddit.uk · 11 hours ago

There are technical solutions to this. You update half your servers, and then if they die you just disconnect them from the network while you fix them and then have your own unaffected servers take up the load. Now yes, this doesn’t get a fixout quickly, but if you update kills your entire system, you’re not going to get the fix out quickly anyway.

floquant@lemmy.dbzer0.com · edit-2 5 hours ago

Congratulations, now your “good” servers are dead from the extra load and you also have a queue of shit to go through once you’re back up, making the problem worse. Running a terabit-scale proxy network isn’t exactly easy, the amount of moving parts interacting with each other is insane. I highly suggest reading some of their postmortems, they’re usually really well written and very informative if you want to learn more about the failures they’ve encountered, the processes to handle them, and their immediate remediations

Cloudfare outage post mortem

Cloudfare outage post mortem

Cloudflare outage on November 18, 2025