Jul 24, 2026
Policy

Microsoft says Azure West US outage stemmed from fiber maintenance bug

Microsoft said routine network maintenance removed too many IP routes, disrupting Azure’s West US region for nearly five hours.

Dominic Okoye

By Dominic Okoye · Staff Writer

· 3 min read

Microsoft says Azure West US outage stemmed from fiber maintenance bug
Photo: The Register

Microsoft said an Azure West US outage on July 23 was triggered by a bug in its maintenance workflow that removed more network routes than intended, disrupting traffic into and out of the California region for nearly five hours. The incident matters for cloud operators because it came from a safety-controlled maintenance process, the kind of internal guardrail customers expect hyperscalers to get right.

According to Microsoft’s preliminary post-incident review, the problem began at 14:44 UTC, or 7:44 a.m. Pacific time, when the company started routine device maintenance. Microsoft said multiple Azure services soon began reporting degradation, though its review did not provide a full list of the affected services.

One minute after the maintenance began, Microsoft said networking teams, service teams and incident responders started looking at traffic anomalies, routing behavior, packet loss signals and recent changes. The company said the incident first appeared as large-scale route churn across its wide-area network. Later investigation tied the route removal to a datacenter in the West US region.

What caused the Azure West US outage?

Microsoft attributed the outage to a bug in a request conversion system used during fiber maintenance. The system wrongly included additional devices in the maintenance event, causing IP routes to be removed from more devices than Microsoft intended.

The company said its standard process for this kind of maintenance is supposed to isolate specific network paths while confirming that at least one of two redundant paths stays healthy. Microsoft also said it runs safety checks intended to confirm the work will not affect customers. In this case, those controls did not prevent the change from affecting live traffic.

Microsoft said the removed routes sat between its datacenter and wide-area network. That meant the failure hit traffic entering or leaving the West US region, rather than being confined to a narrow internal system.

Between 16:00 UTC and 17:45 UTC, Microsoft identified recent fiber maintenance activity and linked it to the routing behavior it was seeing. At 17:45 UTC, the company began rolling back the changes. Microsoft said its wide-area network had returned to normal by 18:26 UTC.

Full recovery took longer. Microsoft said all affected services were fully restored by 19:41 UTC, putting the total elapsed time from the start of maintenance to full service recovery at just under five hours.

The review is preliminary, so it leaves some operational questions unanswered. Microsoft did not say how many customers were affected, which Azure services took the largest hit, or why the conversion bug was able to override the intended maintenance scope. For enterprise users running production systems in a single cloud region, the incident is another reminder that regional redundancy and vendor-maintained network redundancy are different risk controls.

This story draws on original reporting from The Register.

More from Policy

All Policy →