Microsoft's Status Page Said 'Operational.' Outlook Was Down for 12 Hours.

Cover Image for Microsoft's Status Page Said 'Operational.' Outlook Was Down for 12 Hours.
Jagdish Patil
Jagdish Patil

Reports started landing on Downdetector around 8:30 AM Pacific on August 31. By 9:30 AM, they'd peaked near 6,000. Users couldn't send or receive email. Some couldn't even get Outlook to open. One person, unable to reach anyone by email, resorted to physically distributing stamps and envelopes to their team.

For the first several minutes of all this, Microsoft's own 365 status page listed every product as operational. Not degraded. Not investigating. Operational.

A Gap Between What Broke and What Got Reported

Microsoft eventually caught up. The official Microsoft 365 Status account on X acknowledged degraded Exchange Online functionality shortly after the Downdetector spike, and the admin center began surfacing a fuller incident description: users unable to connect to Exchange Online, delayed or failed message delivery, authentication errors, intermittent failures across mailbox operations. But that first window, where thousands of real users were already locked out of email while the public status page said nothing was wrong, is the part worth sitting with.

Microsoft eventually named a preliminary root cause: an issue within a core authentication configuration used by multiple Microsoft 365 services. Not a single service failing in isolation, but a shared authentication layer that Exchange Online, SharePoint, OneDrive, Teams, Purview, and Defender XDR all depend on. When Microsoft tried to fix it, the first mitigation didn't fully take. Engineers spent hours re-examining recent changes, testing whether reverting a recent update would help, and re-applying authentication components across affected infrastructure in stages.

Twelve Hours, Not Twelve Minutes

This wasn't a blip. Reports first appeared around 8:30 AM Pacific. Microsoft didn't call the incident resolved until roughly 8:30 PM Pacific, after a day of partial fixes, renewed testing, and staged recovery across different services. Search functionality inside Exchange stayed broken for a couple of hours even after core mail delivery started working again. For a large chunk of a workday, on one of the most widely used business email platforms in the world, sending and receiving mail was not reliable.

The scope tracked with the outage's origin in shared authentication infrastructure rather than a narrow bug. Downdetector reports came in from across the US, and international reports surfaced too, spanning Latin America and beyond. When the thing that breaks is authentication, everything downstream of it breaks at once, which is exactly what makes this kind of incident slower to fix than a single misbehaving service.

The Actual Lesson

A vendor's status page is written by the vendor, updated on the vendor's timeline, and calibrated to the vendor's own detection systems. None of that is dishonest. It's just structurally slower than the experience of the person who can't send an email right now. Microsoft's status page eventually reflected reality. It just took real user impact, hitting thousands of people and generating public pressure on social media, before that reflection caught up.

This is the same gap we've written about with Cloudflare's and Anthropic's own incidents this year: the provider's account of what happened and your account of what happened can genuinely disagree for the first stretch of an outage, sometimes for the better part of an hour. If your business, or your client's business, depends on that provider working, waiting for their status page to update means you find out last, not first.

The fix isn't distrust of the vendor. Microsoft, Cloudflare, and every major provider generally do get to an accurate public incident report. The fix is not depending on that report to be the first thing that tells you something's wrong. Your own monitoring, checking the specific thing you actually depend on, whether that's a login flow, an API call, or a site that relies on that provider underneath it, will almost always know before the provider's own status page says so.

Waiting for someone else's status page to update means finding out last. SIOPS checks what you actually depend on and fires an alarm that gets through Do Not Disturb, starting free at siops.app.