Google Cloud outage exposes single-site risk in managed services
A 15-hour incident in europe-west4-a hit three Google Cloud services after a power and cooling failure at one datacenter.
By Dominic Okoye · Staff Writer
· 3 min read
Google Cloud suffered a 15-hour outage last week across VMware Engine, NetApp Volumes and Bare Metal Solutions in its europe-west4-a zone after a power issue led to cooling failure at a datacenter. The incident matters for cloud buyers because Google’s own report indicates those three services depended on a discrete facility inside the zone, a finer-grained dependency than many customers can easily see.
Google said in its incident report that the datacenter serving Google Cloud VMware Engine, Bare Metal Solutions and NetApp Volumes for europe-west4-a experienced a power failure, followed by a cooling failure. The company also said an electrical fault on the utility grid upstream of the datacenter disrupted electrical distribution gear and cooling equipment.
Google has not explained how the upstream electrical fault produced that disruption. The company said it turned down workloads to protect customer data from risks associated with operating infrastructure in a high-temperature environment. Google told The Register that its incident analysis is still ongoing and that it would provide more detail when available.
Resilience claims meet facility-level dependencies
The episode is a reminder that a cloud “zone” is not a complete description of the physical architecture behind every service. Google, like other hyperscalers, tells customers to spread workloads across zones and regions to improve resilience. The outage shows that some managed services may still have dependencies on a single datacenter within a zone, even when other parts of that zone continue operating.
The Register said it asked Google whether the affected site had generators or other energy sources, and why workloads had to be turned down if backup power was available. It also asked whether Google tells customers when particular services are tied to a single datacenter. Google had not answered those questions at the time of publication, according to The Register.
Biswajeet Mahapatra, a principal analyst at Forrester, told The Register that customers are often advised to use multiple zones and regions but get limited visibility into whether a managed service has a single-datacenter dependency within a zone. He said some organizations may assume the cloud abstraction includes more facility-level redundancy than specialized services actually have.
Mahapatra also said the architecture is not unusual across hyperscalers. AWS, Azure and Google all run services that depend on dedicated hardware, storage platforms or tightly coupled infrastructure that may not be distributed in the same way as core compute and storage.
A prior Google Cloud outage raised similar questions
Gartner Director Analyst Adrian Wong pointed The Register to a 2023 Google Cloud incident in europe-west9-a, where Google said a water leak originated in a non-Google area of the facility. Google uses Spanner to replicate data across zones, but Wong said the Spanner configuration in that flooded zone did not work once one building became unavailable.
Wong told The Register that it is difficult to determine how an individual cloud region is built, and that Gartner customers are often surprised by that lack of visibility.
Google’s incident report apologized for the disruption and said a final report will describe preventative actions. That may help customers understand Google’s remediation for this outage. It still leaves a broader procurement and architecture issue: customers buying specialized cloud services need to know whether the service’s physical dependencies match the resilience model they believe they are paying for.
This story draws on original reporting from The Register.