Skip to content
theregister·

🚨Google Cloud Outage: Engineer Pulled All Cables

Google Cloud Outage Caused by Engineer Mistake

TL;DR

A Google Cloud outage in September was caused by an engineer disconnecting all fiber-optic cables during maintenance. The incident highlights the importance of robust fail-safes and training.

Google Cloud experienced a partial outage on September 1st in the us-central1-b region when an engineer disconnected all fiber-optic cables during routine maintenance. This caused a 100% traffic flow drop, making VMs unreachable from the outside world. The incident underscores the need for better fail-safes and training to prevent such human errors. The outage lasted from 07:41 to 11:52 PT, affecting network connectivity and resource isolation. Google's resilience provisions moved traffic to healthy capacity elsewhere, but the incident highlights the importance of human error prevention in datacenter operations.

Google Cloud Outage: Engineer Pulled All Cables — theregister

Key Points

1

Google Cloud outage occurred on September 1, 2023, in the us-central1-b region.

2

The outage lasted from 07:41 to 11:52 PT, affecting network connectivity and resource isolation.

3

Traffic flow drop rates for resources in the affected area reached 100% during the peak of the incident.

4

Google's resilience provisions moved traffic to healthy capacity elsewhere in the region.

5

The incident was caused by the physical disconnection of network fiber-optic cables during routine maintenance.

Why It Matters

If you're managing critical infrastructure in a datacenter, this incident highlights the need for robust fail-safes and comprehensive training. The 100% traffic drop and 100% VM unavailability underscores the importance of human error prevention. Google's resilience provisions moved traffic to healthy capacity, but the incident raises questions about the adequacy of current safety measures.

google-cloudoutagedatacenterinfrastructurehuman-error

Frequently Asked Questions

Why does this matter?

If you're managing critical infrastructure in a datacenter, this incident highlights the need for robust fail-safes and comprehensive training. The 100% traffic drop and 100% VM unavailability underscores the importance of human error prevention. Google's resilience provisions moved traffic to healthy capacity, but the incident raises questions about the adequacy of current safety measures.

What happened?

A Google Cloud outage in September was caused by an engineer disconnecting all fiber-optic cables during maintenance. The incident highlights the importance of robust fail-safes and training.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,463 builders reading daily.

Also get