Drooid Logo
Back to story perspectives

Full Breakdown

Major AWS Outage Disrupts Global Internet Services

10/23/2025, 8:48:04 PM

Overview of the Outage

On October 20, 2025, Amazon Web Services (AWS) experienced a significant outage originating from its US-EAST-1 data center in Northern Virginia, which lasted approximately 14 hours. This incident affected over 3,500 companies across more than 60 countries, leading to disruptions in various sectors, including finance, social media, and e-commerce. Downdetector recorded over 16 million user reports, marking this as one of the largest internet outages in history.

Technical Causes of the Outage

The outage was primarily triggered by issues with AWS's DynamoDB service, which experienced increased API error rates due to a "latent defect" in its automated DNS management system. This defect caused endpoint resolution failures, preventing many applications from connecting to the necessary services. As the situation escalated, problems with network load balancers and Elastic Cloud Compute (EC2) instances compounded the issue, leading to widespread service disruptions.

Impact on Services and Users

The outage had immediate and far-reaching effects. Major platforms such as Snapchat, Reddit, and Lloyds Bank were rendered inoperable, while users of smart devices, including Eight Sleep's smart mattresses, reported being stuck in uncomfortable positions due to loss of connectivity. The incident highlighted the internet's reliance on a few dominant cloud providers, with AWS holding approximately 37% of the cloud market.

Criticism and Response

Critics have raised concerns about the systemic risks associated with over-reliance on centralized cloud services. Experts argue that such outages expose vulnerabilities in the current infrastructure, emphasizing the need for businesses to diversify their cloud providers to mitigate risks. AWS has since apologized for the disruption, stating, "We know how critical our services are to our customers, their applications and end users, and their businesses." The company has committed to learning from this incident and improving its infrastructure resilience.

Broader Implications

This outage serves as a wake-up call regarding the concentration of internet services among a few providers. The incident has prompted discussions about the necessity for regulatory oversight and the implementation of multi-cloud strategies to enhance resilience. As businesses increasingly integrate AI and cloud services into their operations, the potential for similar disruptions raises questions about the reliability of these technologies.

What's Next?

In the aftermath of the outage, AWS is undertaking a thorough investigation to prevent future occurrences. The company plans to implement changes to its DNS management processes and enhance its incident response strategies. Additionally, there is growing pressure on policymakers to treat cloud infrastructure as critical components of national resilience, ensuring that such systemic risks are addressed proactively.

Verbatim Quotes

  • “We apologize for the impact this event caused our customers,” — Amazon Web Services
  • “If there’s an outage and you rely on AI to make your decisions and you can’t access it, that’s going to have an effect on performance,” — Tim DeStefano, Associate Research Professor at Georgetown University
  • “Design for failure (because it will happen).” — Lydia Leong, Gartner

Conflicting Reports & Gaps

While AWS has attributed the outage to internal technical failures, some critics speculate that recent layoffs within the company may have contributed to the incident. AWS has denied any connection between the layoffs and the outage, stating that the issues were purely technical.

This incident underscores the critical need for businesses to reassess their reliance on single cloud providers and develop robust contingency plans to ensure operational continuity in the face of potential disruptions.