Skip to main content

On Friday, a major tech outage hit globally, thanks to a faulty software update from the cybersecurity firm CrowdStrike. The glitch brought flights to a standstill, knocked banks and media outlets offline, and disrupted hospitals, small businesses, and other services. This incident underscores just how fragile our tech-dependent world can be, especially when it hinges on just a few key players.

CrowdStrike has clarified that this wasn’t a hack or cyberattack, but an issue with an update affecting computers running Microsoft Windows – they have apologized and are working on a fix. This outage, though, is a clear reminder of our heavy reliance on IT systems and the importance of being prepared for such disruptions.

Software Glitch Shuts Down Image

Key Takeaways

1. Documenting Encryption Recovery Keys

Ensure that BitLocker and other encryption recovery keys are documented on paper or with offline access not reliant on the same or similar infrastructure.

2. Ability to Quickly Roll Back Patches and Updates

Implement point-in-time snapshots to allow for quick rollback of patches and updates.

3. Backups and Disaster Recovery Systems

Maintain backups and disaster recovery systems, including offline or asynchronous backups, for data recovery in case of an outage.

4. Thorough Disaster Recovery Plan

Develop a comprehensive and thorough disaster recovery plan, detailing how to recover unbootable systems remotely when quick physical access is not possible, and estimate the time required to recover from a mass outage impacting over 50% of systems.

5. Flexible/Scalable IT Staffing

Consider outsourced IT support and staff augmentation – flexible and scalable IT staffing can come in handy in these types of instances.

6. Vendor Risk Management

Understand the critical risks associated with each of your primary IT vendors, such as CrowdStrike and Microsoft, and manage those risks effectively.

7. Alternative Vendor Selection

Identify fail-back vendor options to ensure continuity if a main vendor fails.

8. 24×7 Monitoring and Response

Implement continuous monitoring and response capabilities to detect and address issues immediately.

9. Careful Management of Patching

Avoid deploying patches to all systems all at once to prevent widespread outages.

10. Warm Standby Failover/Recovery Environments

Maintain warm standby failover/recovery environments that are not patched on the same cycle as production environments.

The Bottom Line

While the recent incident with CrowdStrike was not a security-related event, it teaches us important lessons about the possibility of an intentional large-scale outage caused by a deliberate security event. We can use these lessons to prepare for such an attack, as well as other unintentional outages like this one, to be better ready for the future.

New Charter operating companies are well-equipped to help clients implement and manage these critical measures.