🚨Google Cloud Outage: Engineer Pulled All Cables
Google Cloud Outage Caused by Engineer Mistake
TL;DR
A Google Cloud outage in September was caused by an engineer disconnecting all fiber-optic cables during maintenance. The incident highlights the importance of robust fail-safes and training.
Google Cloud experienced a partial outage on September 1st in the us-central1-b region when an engineer disconnected all fiber-optic cables during routine maintenance. This caused a 100% traffic flow drop, making VMs unreachable from the outside world. The incident underscores the need for better fail-safes and training to prevent such human errors. The outage lasted from 07:41 to 11:52 PT, affecting network connectivity and resource isolation. Google's resilience provisions moved traffic to healthy capacity elsewhere, but the incident highlights the importance of human error prevention in datacenter operations.

Key Points
Google Cloud outage occurred on September 1, 2023, in the us-central1-b region.
The outage lasted from 07:41 to 11:52 PT, affecting network connectivity and resource isolation.
Traffic flow drop rates for resources in the affected area reached 100% during the peak of the incident.
Google's resilience provisions moved traffic to healthy capacity elsewhere in the region.
The incident was caused by the physical disconnection of network fiber-optic cables during routine maintenance.
Why It Matters
If you're managing critical infrastructure in a datacenter, this incident highlights the need for robust fail-safes and comprehensive training. The 100% traffic drop and 100% VM unavailability underscores the importance of human error prevention. Google's resilience provisions moved traffic to healthy capacity, but the incident raises questions about the adequacy of current safety measures.
Frequently Asked Questions
Why does this matter?
If you're managing critical infrastructure in a datacenter, this incident highlights the need for robust fail-safes and comprehensive training. The 100% traffic drop and 100% VM unavailability underscores the importance of human error prevention. Google's resilience provisions moved traffic to healthy capacity, but the incident raises questions about the adequacy of current safety measures.
What happened?
A Google Cloud outage in September was caused by an engineer disconnecting all fiber-optic cables during maintenance. The incident highlights the importance of robust fail-safes and training.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,463 builders reading daily.