Blog · Operations

A file that doubled in size took Cloudflare down for hours. It wasn't an attack

A permissions change in a database caused a configuration file to double in size and exceed an internal limit. The file spread across the entire network, and traffic started failing at 11:20 UTC; it took until 17:06 for everything to return to normal. At first the team thought it was an attack. It was a file generated by Cloudflare's own systems.

Network switch with cables
Network switch with cables. Cropped to 16:9. Photo: ProjectManhattan · CC BY-SA 3.0 · Wikimedia Commons

What happened, hour by hour

At 11:05 UTC on November 18, 2025, Cloudflare applied a change to the access permissions of a ClickHouse cluster. A query that builds the feature file for the bot management system started returning duplicate rows, and the file doubled in size. That file propagated to every machine on the network.

The software that reads it has a limit of 200 features. When the file exceeded that limit, the Rust code called unwrap on an error and panicked, and every request ended in a 5xx error. The CDN and security services failed, along with Turnstile, Workers KV, Access, and dashboard login.

At 13:05 the team applied a bypass for Workers KV and Access. By 14:30 core traffic was flowing almost normally, after a correct version of the file was distributed. By 17:06 all services had recovered. Cloudflare described it as its worst outage since 2019.

Why did it look like an attack?

The file was generated every five minutes, and the ClickHouse cluster was being updated in parts. If the query ran on a part that had already been updated, it produced a bad file; if not, a good one. So the network kept going down and recovering. On top of that, Cloudflare's status page, hosted outside its infrastructure, stopped working by coincidence. The team feared it was a continuation of the Aisuru DDoS attacks.

The report is clear: the problem was not caused, directly or indirectly, by an attack or by malicious activity of any kind.

The general lesson

The failure came from inside, from a file the company itself generated and that its systems trusted without validating. Among the measures Cloudflare announced are validating those files as if they were user input, adding global kill switches to turn off features, and reviewing how each proxy module fails.

Sources

  1. Cloudflare, official incident report

Consulting: technical leadership →

← Back to the blog