Resilience at Cloud Scale: Azure CTO on Outages, Hardware, and AI
70% of cloud outages are change-related; Azure now uses AI/ML and chaos engineering to detect and resolve failures without humans.
“70% of outages in the cloud and kind of at the industry scale are change related in some way.”
Azure's CTO recounts a 2014 storage outage that led to the company's safe deployment policy, and explains how Azure standardized health measurement via SLIs and now applies ML to define service health and shrink time-to-resolution by removing humans from the loop. It's a substantive engineering account of cloud reliability and emerging AI-driven operations, but light on broad AI-industry signal.