Blog · Cloud

A cooling failure in Virginia left Coinbase unable to operate for hours

On May 8, 2026, cooling failed at an AWS data center in us-east-1, in Northern Virginia, and Coinbase went offline for several hours. In October 2025, a DNS failure in the same region had affected 64 internal AWS services and hundreds of applications. If your whole system lives in a single availability zone, your availability depends on what happens in those buildings.

Virginia Tech data center (illustrative image)
Virginia Tech data center (illustrative image). Cropped to 16:9. Photo: Christopher Bowns · CC BY-SA 2.0 · Wikimedia Commons

What happened on May 8

AWS reported elevated internal temperatures at one of its data centers in us-east-1 due to a failure in the cooling system. It traced the problem to a single availability zone and shifted traffic to the other zones in the region. Additional cooling came online a couple of hours after the first reports, but a later update warned that safely restarting all affected systems was taking longer than expected.

Coinbase confirmed that its problems stemmed directly from the AWS event and, after several hours with degraded markets, reported that everything was operating again. AWS's recommendation during the incident was the usual one: customers in the affected zone had to fail over to another zone.

The same region, seven months earlier

On October 20, 2025, a DNS problem in us-east-1 affected 64 internal AWS services. Snapchat, Reddit, Zoom, Venmo, and Robinhood had problems, and so did Coinbase. Downdetector received more than 6.5 million reports about more than 1,000 companies.

The two failures are different. The May failure stayed within one zone: anyone whose system was spread across several zones could keep going. The October failure hit a region-wide service: there, spreading across zones isn't enough, and what protects you is having your critical systems ready to run in another region.

What an availability zone covers

A zone is a group of one or more data centers with power, cooling, and networking independent from the other zones in the region. It is designed so that a failure in one zone doesn't drag down the others. Failover works if you were already there: database replicas in another zone, enough capacity to absorb the traffic, and a switchover tested in advance. If not, AWS's recommendation comes too late.

Sources

  1. WWWhat's new, 5/9/2026
  2. BBC Mundo, 10/20/2025
  3. ABC7, 10/20/2025

Consulting: AWS architecture review →

← Back to the blog