|

Three Data Centers, One Failure Domain: The Cloudflare Outage

The Grey Rhino Files cover with a rhinoceros and episode number 002

On November 2, 2023, one of Cloudflare’s three core data centers near Portland started having power problems. That is already a bad day for an infrastructure team, but it was not an unplanned failure mode. For years, the company had been building its control plane so it could survive the complete loss of one facility. The three sites were physically separated, critical systems were gradually moving into an active-active high-availability cluster, and the remaining two sites were expected to keep running if the third disappeared completely.

That is why the interesting part of this story is not that a large data center lost power. Cloudflare already considered that an acceptable failure mode. A large part of the system did survive exactly as designed. The problems started where highly available services depended on systems that were not highly available.

Physically, Cloudflare had three data centers. For part of the software dependency graph, the system still effectively had one failure domain.

The Data Center the Architecture Was Supposed to Lose

Cloudflare referred to the largest of the three Oregon sites as PDX-04. It was operated by Flexential and hosted the company’s largest analytics cluster, more than a third of the machines in the main HA cluster, and services that had not yet been moved to high-availability infrastructure. That unevenness was part of the platform’s normal evolution: most critical control-plane systems already ran in HA, while newer products could move there over time.

Cloudflare also did not try to give every service the same failure model. Logging and analytics, for example, could degrade differently from the control plane. Logs were produced across Cloudflare’s global distributed network and could wait for a while until the central infrastructure returned. That is a reasonable engineering trade-off: delayed analytics and the inability to change customer configuration are different kinds of degradation, and there is no reason to pay the same availability cost for both.

At about 08:50 UTC, Portland General Electric was doing unplanned work that took one of the facility’s independent power feeds offline. Flexential started its generators. Around 11:40, a ground fault occurred on a high-voltage transformer. Cloudflare said a connection to the earlier PGE work was possible, but it explicitly noted that the causal link was not confirmed. The electrical protection system then operated in a way that stopped all ten generators from supplying power along with the external feeds.

The remaining layer was the UPS battery system, designed to provide about ten minutes of power while generators start. From what Cloudflare could see on its own equipment, servers began shutting down after roughly four minutes. By 12:01, much of the data center was unavailable.

This was a serious physical failure, but it was also exactly the boundary the software architecture was supposed to survive. From Cloudflare’s point of view, the exact reason the building stopped delivering power did not really matter. The architectural guarantee was simpler: one of the three core sites could disappear completely, and the critical part of the control plane should continue running on the other two.

Mostly, it did. Not all of it.

High Availability Until the First Single-Site Dependency

After PDX-04 disappeared, the first layer looked correct. The other two data centers were still online, the HA cluster was still running, and the network and security data plane continued to process customer traffic. Some analytics and less critical services degraded as expected. But several systems that were formally inside the high-availability cluster also became unavailable.

The reason was dependencies. Cloudflare specifically called out Kafka and ClickHouse, two critical services used for log processing and analytics. They ran only in PDX-04, while other services inside the HA cluster depended on them. Cloudflare described those relationships as too tight and said they should have degraded more gracefully.

A simplified view looks like this:

A distributed service with a single-site dependencyService replicas run in PDX-04, Site B, and Site C, but all depend on Kafka, which exists only in PDX-04.ServicePDX-04Service replicaSite BService replicaSite CService replicaKafkaonly in PDX-04

The service itself looks fine: copies run in several data centers, so two remain alive when one site disappears. But the service’s availability is not determined only by its own deployment. If it cannot perform a critical function without Kafka, and Kafka exists only in PDX-04, that dependency brings the same data center back into the failure domain.

That makes the incident more interesting than a simple story about a hidden dependency. High availability is easy to inspect locally: how many replicas exist, which zones they run in, whether a primary can fail over. It is much harder to answer the same question for the full critical path. A service can be physically distributed and still remain logically tied to one site because of a database, queue, configuration service, or another transitive dependency.

Cloudflare was not just trusting the HA architecture on paper. The company had tested it. It had completely shut down each of the other two core data centers and had even tested losing both at once. It had also tested the HA part of PDX-04. What it had not tested before was the complete loss of all of PDX-04.

That sounds like a small difference in test wording, but it is a large difference in system behavior. In one case, the HA nodes inside PDX-04 disappear while local services at the site remain available. In the other, the entire data center disappears along with any local dependencies. The real outage tested the second case.

Another property of a large, constantly evolving platform appeared at the same time. Cloudflare let product teams move quickly to an alpha stage and then migrate their backends to the standard HA infrastructure over time. Completing that migration was not a formal requirement before a product reached Generally Available (GA) status. So GA described product maturity from the customer’s point of view, but it did not guarantee the same level of infrastructure maturity behind every product.

That approach does not look absurd on its own. Full HA architecture costs time and capacity, while a new product may still change radically or fail to find a market. The incident simply exposed an extra consequence of that gradual path: reliability stops being only a property of the platform architecture and also depends on how far each service has progressed through its migration.

Disaster Recovery Meets Real Traffic

At 12:48 UTC, Flexential managed to start generators again and began restoring power in stages. When the process reached Cloudflare’s lines, some circuit breakers turned out to be faulty. With the recovery timeline still unclear, Cloudflare decided at 13:40 to move the required parts of the control plane to disaster-recovery (DR) sites in Europe. The first services started there at 13:43.

That recovery exposed another property of outages that is easy to underestimate during capacity planning. API requests that had been failing arrived in enough volume to overload the restored infrastructure once the DR services became available. Cloudflare described this as a thundering herd and introduced rate limits to stabilize request volume. By 17:57, the migrated services were stable enough that most customers were no longer directly affected by the incident.

The postmortem does not support inventing a specific retry algorithm or a more elaborate story about queues and reconnects. The observed effect is already interesting: recovery does not necessarily receive the same workload as normal operation. When a service returns after hours of errors, there may already be accumulated demand waiting for it. Capacity that is enough for steady state does not automatically guarantee a calm return after an outage.

Meanwhile, the data center itself was being repaired. At 22:48, Flexential restored both external feeds and confirmed stable power. By then, the Cloudflare team had been working the incident all day. Matthew Prince decided not to begin a complex full bootstrap of the large data center during the night and let most of the team sleep, deliberately accepting a slower recovery to reduce the operational risk of mistakes from exhausted people.

The next morning, PDX-04 had to be brought back almost like a new site. Because several power cycles may have occurred, the state of the systems was treated as unknown, so Cloudflare went through a full bootstrap process. Restoring the configuration-management servers alone took around three hours. After that, thousands of machines had to come back, with individual servers taking anywhere from about ten minutes to two hours and some services needing to start in dependency order.

Cloudflare closed the incident on November 4 at 04:25 UTC. The data plane for network and security services continued to run throughout, but the control plane had major disruptions, raw log services were unavailable for most customers for much of the incident, and some analytics datasets ended up with gaps.

Afterward, Cloudflare launched an internal program called Code Orange and redirected non-critical engineering work toward control-plane reliability for several months. Among the planned changes, Cloudflare said GA products that depended on core data centers should use the HA cluster without software dependencies on a specific site; disaster recovery for GA should be tested; and chaos testing should include removing each core data center completely.

For the story itself, the remediation list is less interesting than what happened next. Five months later, the same data center lost power again.

The Same Failure, a Different System

On March 26, 2024, the Flexential facility again suffered a complete power loss for Cloudflare infrastructure. This time, the physical cause was different: according to Flexential, four switchboards serving Cloudflare’s cages shut down at the same time. For the software, though, it was almost a perfect repeat of the earlier experiment: one of the three Portland data centers simply disappeared.

By then, Cloudflare had spent several months changing the architecture and had run an internal cut test in February to verify automatic failover. When the facility lost power at 14:58 UTC on March 26, systems began failing over automatically. By 15:05, the API and Dashboard were operating normally again without human intervention. In November, many of the same control-plane products had seen at least six hours of downtime; in March, they recovered within minutes, and some had no noticeable impact during failover.

One of the changes was moving configuration databases into a high-availability topology and pre-provisioning enough spare capacity for the other two data centers to accept the load. During the March incident, more than one hundred databases across more than twenty clusters automatically failed over away from the affected site. Cloudflare described this database failover as the result of more than a year of work, not only the five months of Code Orange. The company also said this capability was now tested weekly.

Not everything became highly available in five months. Analytics still depended on the data center and recovered later that day. Cloudflare described this as the expected state at the time because the analytics migration was still in progress. That detail makes the comparison between the two incidents more useful: there is no need to pretend the first outage made the whole system perfect. The changes had a clear boundary, and the second outage quickly showed which ones were already working.

Even the data center’s cold start changed. Cloudflare later estimated the full cold-start process at roughly 72 hours in November; in March, it took around ten hours. This is less a story about a team “fixing mistakes” and more about what happens when a failure passes through the full dependency graph once and turns invisible assumptions into things that can be tested directly.

Before November, Cloudflare already had three data centers, active-active clusters, disaster recovery, and real failure tests. All of those things were real. And yet there was still a dependency graph between “the service has replicas across several data centers” and “the service survives the loss of a data center.”

That was the real boundary of redundancy: not the last replica visible on a deployment diagram, but the first critical dependency that still shared the same failure domain.

Sources