Amazon Web Services (AWS) has published 18 post-event summaries for major outages between April 2011 and October 2025, and 12 of those outages were in its US East (N. Virginia) region. The most recent, in October 2025, disrupted DynamoDB and much of that region for about 14.5 hours after a fault in automated DNS management. The longest, in April 2011, left part of one zone's storage unusable for almost three days.

This page lists every one of those outages with its date, duration, and root cause, taken from AWS's own summaries. It also covers three large incidents from 2026 that have no summary yet, and what site reliability engineers (SREs) can learn from the pattern. It was last checked against AWS's index on October 11, 2026.

Dot chart of AWS's 18 documented major outages by year from 2011 to 2025. Twelve were in US East (N. Virginia) and six were in other regions.
Each dot is one outage that AWS documented with a public post-event summary. Twelve of the 18 were in US East (N. Virginia).

Is AWS down right now?

This page is a history, so it cannot tell you about a live incident. For the current state of every AWS service, open the AWS Health Dashboard. If you came here during an outage, the sections on causes and lessons below explain what is probably happening and what tends to recover first.

What counts as a major AWS outage?

This page uses AWS's own bar. AWS publishes a post-event summary when an issue "has broad and significant customer impact". Its index of summaries names three triggers: failure of a significant share of control plane API calls, impact on a significant share of a service's infrastructure, or a total power failure or significant network failure. Summaries stay online for at least five years.

Smaller incidents get a note on the Health Dashboard and no summary. That means the real number of AWS disruptions is higher than 18. The list below is the set that AWS itself judged serious enough to explain in public.

What is the full list of major AWS outages?

Here are all 18, newest first. Each date links to AWS's summary. Durations are worked out from the times AWS gives, and they describe the main period of customer impact.

Every major AWS outage with a post-event summary, 2011 to 2025
When and whereWhat failed, and for how longRoot cause
Oct 19–20, 2025
US East (N. Virginia)
DynamoDB, then EC2 launches, load balancers, Lambda, and many dependent services. About 14.5 hours.A race condition in DynamoDB's automated DNS management left the regional endpoint with an empty DNS record.
Jul 30, 2024
US East (N. Virginia)
Kinesis Data Streams, which disrupted CloudWatch Logs, Data Firehose, ECS, Lambda, and others. About 7 hours.A Kinesis management system mishandled a very large number of low-traffic shards and treated healthy hosts as failed.
Jun 13, 2023
US East (N. Virginia)
Lambda, plus sign-in, EventBridge, and the console. Just under 4 hours.Scaling pushed part of Lambda past a size it had never reached, which exposed a hidden defect in how capacity was allocated.
Dec 7, 2021
US East (N. Virginia)
Internal network congestion that hit EC2 APIs, the console, and many services. About 7 hours, with some effects for 11.An automated scaling activity set off a surge of connections that overloaded internal network devices. A hidden defect stopped clients from backing off.
Sep 2, 2021
Tokyo
Direct Connect links into the region. About 6 hours.A hidden defect in network device software, triggered by a rare pattern of packets.
Nov 25, 2020
US East (N. Virginia)
Kinesis, which disrupted Cognito, CloudWatch, Lambda, and EventBridge. About 17 hours.Adding capacity pushed every front-end server past an operating system limit on threads.
Aug 23, 2019
Tokyo
EC2 and EBS in part of one zone. About 6 hours for most resources.A data centre cooling control system failed, and a bug in its third-party logic blocked the fallback, so servers overheated.
Nov 22, 2018
Seoul
DNS lookups from EC2 instances. 1 hour 24 minutes.A configuration update removed the setting for the minimum number of healthy DNS servers.
Feb 28, 2017
US East (N. Virginia)
S3, plus the many services that depend on it. A little over 4 hours.A mistyped command removed far more servers than intended, including ones that S3's index needed.
Jun 4, 2016
Sydney
EC2 and EBS in one zone. Power was out for about 80 minutes, and nearly all instances were back about 9.5 hours after the start.Utility power failed in severe weather and the backup power system did not switch over as designed.
Sep 20, 2015
US East (N. Virginia)
DynamoDB, plus SQS, Auto Scaling, and CloudWatch. About 5 hours.A brief network disruption made storage servers request membership data all at once, and the metadata service could not keep up.
Jun 13, 2014
US East (N. Virginia)
SimpleDB. About 2 hours fully down, then about 2 more with errors.A power loss took storage nodes offline, and a timeout that was set too low made healthy nodes remove themselves.
Dec 17, 2013
São Paulo
One zone lost power, and both zones had about 20 minutes of degraded networking. AWS gives no total duration.A fault at the utility's substation, followed by backup generators failing during the switch.
Dec 24–25, 2012
US East (N. Virginia)
Elastic Load Balancing. Almost 24 hours.A maintenance process was run against production by mistake and deleted load balancer state data.
Oct 22, 2012
US East (N. Virginia)
EBS in one zone, plus EC2 APIs, RDS, and load balancers. About 6 hours for storage and longer for load balancers.A memory leak in a monitoring agent on storage servers, set off when a DNS change failed to reach them.
Jun 29–30, 2012
US East (N. Virginia)
EC2, EBS, RDS, and load balancers in one zone. About 20 minutes without power, then hours of recovery.An electrical storm cut utility power, and backup generators failed to provide stable voltage.
Aug 7, 2011
EU West (Ireland)
EC2, EBS, and RDS in one zone. Power and network returned in about 3 hours, and some storage recovery took 3 days.A utility transformer failed and a controller did not bring the backup generators online.
Apr 21–24, 2011
US East (N. Virginia)
EBS and RDS in one zone, with errors across the region. Almost 3 days until storage in the zone was fully usable.A network change was carried out incorrectly, which cut storage nodes off and set off a storm of data re-copying.

One correction to AWS's own index: it dates the EU West summary August 7, 2014. The outage happened on August 7, 2011, as news reports from that week show.

What AWS outages have happened in 2026?

Three large incidents in 2026 had no post-event summary on AWS's index when this page was last checked. Their details come from AWS's status messages as reported by others, so treat them as less settled than the table above.

Large AWS incidents in 2026 without a post-event summary
When and whereWhat failedReported cause
Mar 1, 2026
UAE and Bahrain
Two of the three zones in the UAE region were significantly impaired, and one facility in Bahrain was affected. AWS advised customers to move workloads to other regions.Data centres were struck during military attacks in the Gulf. AWS said a facility "was impacted by objects that struck the data center", causing a fire.
May 7–8, 2026
US East (N. Virginia)
EC2 and EBS in one zone. Coinbase halted nearly all trading for about 8 hours, and a monitoring firm counted more than 150 affected cloud services.Several cooling units failed in one data hall, and servers shut down to protect themselves from heat.
Jul 24, 2026
US West (Oregon)
Connections into the region for about an hour, across ten services including EC2, load balancers, and API Gateway.Networking hardware that carries routes between the region and the Seattle area.

Sources: Data Center Dynamics and InfoQ for March, Coinbase's own postmortem and StatusGator for May, and IncidentHub for July.

What was the biggest AWS outage?

It depends on whether you measure reach or length. The October 2025 outage reached furthest, and the April 2011 outage lasted longest. These five are the ones engineers still refer to:

  • October 2025, DynamoDB DNS. Two automated processes that manage DynamoDB's DNS records raced each other. One applied an old plan and the other then deleted it, which left the region's main DynamoDB address with no IP addresses. DynamoDB recovered in about three hours, but EC2 could not launch new instances for another eight, because its own systems depended on DynamoDB.
  • December 2021, internal network. A scaling activity caused many internal clients to open connections at once. The devices linking AWS's internal network to its main network became congested, and retries kept them congested. AWS's own monitoring ran on the affected network, so engineers were working with limited visibility.
  • November 2020, Kinesis. A small addition of capacity made every front-end server open more threads than the operating system allowed. Kinesis sits underneath CloudWatch and Cognito, so monitoring and sign-in failed with it.
  • February 2017, S3. An engineer following a runbook typed one input to a command incorrectly. It removed servers from two core S3 subsystems, which then needed a full restart that took hours because they had not been restarted in years.
  • April 2011, EBS. A routine network change sent traffic to the wrong network. Storage nodes lost contact with their replicas and all tried to create new copies at once. The zone ran out of space, and recovery took days.

Why does US East (N. Virginia) have so many outages?

US East (N. Virginia), which AWS calls us-east-1, accounts for 12 of the 18 documented outages. Two reasons explain most of that.

First, it is AWS's oldest region and one of its largest, so more services and more customers are exposed to anything that goes wrong there. Some of the outages began when a service reached a size that it had never reached before. The 2020 and 2023 summaries each describe a limit that was hit for the first time in this region.

Second, some AWS-wide functions depend on it. The summaries show this directly. In December 2021, Route 53's management APIs were impaired. In October 2025, customers in other regions had trouble signing in to the console with certain credentials, and some Redshift features failed in every region.

The practical point for an engineer is that running in another region does not fully separate you from us-east-1. Find out which of your dependencies have a control plane there, and test what happens when it stops answering.

What causes AWS outages?

Hidden software defects and capacity limits cause the most, followed by power or cooling failures, then mistakes during operations. The chart groups the 18 documented outages by the first thing that went wrong.

Bar chart of what triggered AWS's 18 documented major outages: 8 by a latent software defect or capacity limit, 6 by a power or cooling failure, and 4 by an operator or change error.
This grouping is ours, based on the first trigger each summary names. Several outages had more than one contributing cause.

Four patterns repeat across the summaries:

  • Routine work is the usual trigger. Adding capacity (2020), automatic scaling (2021 and 2023), a single command (2017), and a configuration update (2018) each started a major outage. None of them was an unusual or risky project.
  • Retries make a bad situation worse. In 2011, 2015, and 2021, clients that kept retrying overloaded a system that was already struggling. AWS's fixes each time included making clients wait longer between attempts.
  • Monitoring fails along with the thing it monitors. In 2017 the AWS status page could not be updated because its tooling depended on S3. In 2020 and 2021 AWS's own dashboards and support tools were affected too.
  • Backup power and cooling fail when they are needed. In 2011, 2012, 2013, 2016, and 2019, the outage happened because a generator, a breaker, or a control system did not take over as designed.

What can SREs learn from AWS outages?

The same six lessons appear again and again in AWS's summaries, and each one applies to a system of any size:

  1. Survive the loss of one zone, and prove it. After the May 2026 outage, Coinbase wrote that "the loss of an entire zone should result in reduced capacity, not unavailability". A design that should survive is not the same as one that has been tested. Our reliability testing lesson shows how to run that test safely.
  2. Do not treat several zones as disaster recovery. Zones in one region share a city, a power grid, and a control plane. In March 2026 AWS told customers in the Gulf to move to other regions. Decide how much data you can afford to lose and how long you can be down, which the data recovery lesson calls RPO and RTO.
  3. Make clients back off. A retry is a new request. When thousands of clients retry at once, they can hold a recovering system down. The capacity and overload lesson explains how this turns one failure into a wider one.
  4. Keep your monitoring independent. If your dashboards and your status page run on the system they watch, you lose them exactly when you need them. See the monitoring lesson.
  5. Slow down dangerous commands. After 2017, AWS changed its tool to remove capacity more slowly and to refuse a removal that would take a system below its safe minimum. Any command that can remove capacity deserves the same two guards. The safe releases lesson covers limits like these.
  6. Write it down afterwards. AWS's summaries are useful because they name the trigger, the contributing causes, and the specific changes that followed. That is the structure of a good blameless postmortem.

How often does AWS have a major outage?

AWS has published 18 summaries for the 14 and a half years from April 2011 to October 2025, which is a little more than one a year. They are not evenly spread. There were three in 2012, two in 2011 and in 2021, and none in 2022.

That rate describes outages serious enough for a public summary. Shorter and narrower incidents are more frequent, and they appear only on the Health Dashboard, which keeps 12 months of history.

Exercise: find your single points of failure

A small company runs its booking service on AWS. Everything is in one zone of US East (N. Virginia): two web servers, one database with nightly backups stored in the same region, and a status page hosted on the same servers. List what breaks if that zone loses cooling for eight hours, then decide what to change first. Work out your own answer before comparing your reasoning with this one:

Everything breaks at once. The web servers and the database are in the failed zone, so the service is down for the full eight hours. The status page is down too, so customers get no explanation. The backups survive, but restoring them needs a working zone and someone who has practised the restore.

The first change is to run the web servers and a database replica in a second zone, then switch one zone off on purpose to check that the service keeps working. The second is to move the status page somewhere that does not depend on these servers. The third is to copy backups to another region and time a restore, because a zone is not the largest thing that can fail.

Quick answers

When was the last major AWS outage?

The most recent outage with a post-event summary from AWS was on October 19 and 20, 2025, in US East (N. Virginia). Since then, large incidents were reported in March, May, and July 2026, and AWS had published no summary for them by October 11, 2026.

What caused the October 2025 AWS outage?

A race condition in the automation that manages DynamoDB's DNS records. It left the service's regional address with no IP addresses, so nothing could reach DynamoDB in US East (N. Virginia). Services that depend on DynamoDB, including EC2 instance launches, then failed for hours longer.

What caused the 2017 S3 outage?

A mistyped command. An engineer meant to remove a small number of servers from a billing system and removed a much larger set, including servers that two core S3 subsystems needed. Restarting those subsystems took a little over four hours.

How long did the December 2021 AWS outage last?

About seven hours for the main network congestion, from 7:30 AM to 2:22 PM Pacific time on December 7, 2021. Some services took longer, and the last reported effects ended at 6:40 PM.

Does AWS publish postmortems?

Yes, for its largest outages. AWS calls them post-event summaries and keeps each one online for at least five years. Smaller incidents are recorded only on the AWS Health Dashboard.