Twenty-four hours have passed. Many IT and security departments have canceled their vacations, and the world is trying to recover from what is likely the biggest IT "blackout" in history. We've also noticed confusion in the reports—some blaming Microsoft, others blaming CrowdStrike, and others pointing fingers at both companies. Let's see.
What Happened?
On July 19, CrowdStrike, a cybersecurity company based in Austin (the capital of Texas), released an update for its Falcon platform. Falcon's role is to install and maintain sensors on "mission-critical" terminals running Microsoft Windows to "detect and prevent" threats.
Okay, That Sounds Nice, But...
...the update had a bug. The wave of blue screens began in Australia and "moved west" as the day progressed.
Who Was Affected?
Banks, hospitals, airports, hotels, ground transportation, food services, dozens of government agencies, television channels, e-commerce portals, and many more, including remote workers with assigned laptops. More than 5,000 scheduled flights had to be canceled. Medical visits, consultations, and many surgeries were put on hold. Even 911 and the Olympic Games organization were impacted.
Was It a Cyberattack?
No, and this must be emphasized: it was a bug in the Falcon update distributed by CrowdStrike.
What About Mac and Linux? Android? iOS?
Immune. The bug affects Windows computers that use CrowdStrike Falcon.
How Do You Fix It?
This is where things get complicated. Fixing the error on an isolated machine isn't that hard, but we're talking about organizations with thousands of systems. The first recommendation, although it might seem like a joke, is to restart the machine dozens of times (details below). The second is more technical: manually delete the faulty update files, named C-00000291*.sys (note the asterisk), in the path Windows\System32\Drivers\CrowdStrike.
Restart Dozens of Times? Are You Kidding?
Actually, no. The systems are caught in a boot loop, but at the same time there's a kind of "race" between the network stack and the faulty update. If on one of the restarts the network stack wins, CrowdStrike can install the correct update and allow the system to escape the loop. If the buggy update wins, BSOD.
What If the Machine Uses BitLocker?
Pain and suffering. If the recovery keys to read the affected drives aren't available, the process is essentially a self-inflicted ransomware attack.
Alternatives?
Another recommendation from Microsoft and CrowdStrike is to restore a backup made before July 19, 04:09 UTC, the time when CrowdStrike began distributing its faulty update.
Some Say It Was Microsoft. Why?
Because on Thursday night, several Azure-associated services also suffered an outage across the central United States. The failure extended from July 18 (21:56 UTC) to July 19 (12:15 UTC). "God does not play dice..."
In the End, How Big Was It?
Troy Hunt of Have I Been Pwned? said it was "the biggest IT failure in history", and added that this was basically what we worried about with Y2K, with the difference that it happened now.
So... Will It Happen Again?
If we keep doing stupid things with buggy updates at Ring 0, definitely.
Sources: CrowdStrike, Ars Technica, The Verge, TechRadar, Microsoft