#THE GREY RHINO FILES
The Obvious Failures We Chose to Ignore
What is a “Grey Rhino” in Tech?
Unlike a Black Swan — an unpredictable, highly improbable event — a Grey Rhino is a massive, high-impact threat charging right at us. It’s the outdated database version everyone knew needed an upgrade, the unverified backup strategy, or the missing safeguard on a production cluster.
We all see it coming. We all know it’s dangerous. Yet, we ignore it until it unexpectedly tramples production.
In this series, we unpack real-world IT disasters to understand why these obvious risks were missed and how to stop the next Grey Rhino in its tracks.
- Three Data Centers, One Failure Domain: The Cloudflare Outage
Cloudflare designed its control plane to survive the loss of a data center. A real outage showed that redundancy ends where a critical dependency still shares the same failure domain.
- A Backup You Haven't Restored Is Just a Hypothesis: The GitLab Database Outage
How GitLab lost its production database, discovered that several backup systems were not actually usable, and was saved by a staging snapshot made six hours earlier.