AWS Large-Scale Outage Raises Questions about IT Resilience
On Monday, Amazon Web Services (AWS) experienced a large-scale outage that affected thousands of customers and multiple digital services. Despite mitigation measures taken by the company, services remained in a degraded state. Analysts noted that the incident once again reminded CIOs of the importance of multi-cloud and on-premises hybrid architectures, as well as recovery plans independent of the primary cloud provider.

On Monday morning, Amazon Web Services (AWS) experienced a large-scale outage that affected thousands of customers and triggered cascading failures in multiple digital services. Although initial recovery measures partially alleviated operational difficulties for hundreds of AWS services in the US-East-1 region, the issue was not fully resolved, and Amazon was still locating and fixing the root cause later that day.
According to updates on the company's status page, the Amazon subsidiary providing cloud services attributed the issue to an internal subsystem responsible for monitoring the health of its network load balancers.
"We have taken additional mitigation measures to help recover the underlying internal subsystem responsible for monitoring the health of network load balancers, and we are now seeing recovery in connectivity and APIs for AWS services," the company said around noon Eastern Time on Monday, but still listed service status as "degraded."
Cloud outages can ripple through digital services, disrupting multiple applications at once and hindering business continuity plans. When the affected provider is AWS, which leads its peers in market share, the impact can be further amplified.
According to Gartner estimates, Amazon Web Services attracted 37.7% of all infrastructure-as-a-service (IaaS) spending last year, while Microsoft held a 23.9% market share. Google controlled only 9% of spending last year.
John Annand, practice lead for digital infrastructure at Info-Tech Research Group, said cloud outages serve as a wake-up call for chief information officers (CIOs), helping them assess the resilience of their IT assets.
"Trying to reduce any risk to zero comes with exponentially rising costs," Annand said. "The lower you want the risk, the more it costs."
IT stress testing
Annand noted that vendor selection is only part of the resilience puzzle for CIOs. But from an architectural perspective, cloud systems that rely on overlapping vendors can become overly complex.
"It looks good on paper, people talk about it at meetings, but they don't actually do it," Annand said. "You have to weigh the effectiveness and ease of use of cloud platforms, and then try to plan for when you know outages might happen."
Roy Illsley, chief analyst for IT operations at Omdia, said the key takeaway from such outages for CIOs is to develop a dual-source strategy.
"This event shows that even companies like AWS can be affected, and unless you have a contingency plan, you're in trouble," he said in an email to CIO Dive.
Illsley believes that multi-cloud provides an additional layer of resilience, but migrating workloads between clouds is challenging. Ideally, CIOs should consider combining multi-cloud with on-premises environments, but he also cautioned that this strategy is more costly and complex to implement.
"There's no silver bullet," Illsley said. "But CIOs must do their due diligence and consider robust recovery plans that are independent of the primary cloud provider."
IT outages can cause significant losses for businesses. According to data released by New Relic last month, operational downtime caused by technical issues costs enterprises a median of $2 million per hour. The company found that cloud service failures are one of the leading causes of IT downtime.
Last year, a faulty CrowdStrike update pushed to Windows devices triggered a massive outage that affected global IT systems. The July 2024 incident resulted in an estimated more than $5 billion in direct financial losses for Fortune 500 companies, with the healthcare industry suffering the largest financial impact.
Analysts and experts previously told CIO Dive that unplanned IT failures can provide an opportunity to reassess business continuity plans.
"The question isn't whether services will go down," Annand said. "It's when they will go down. As a CIO, your job is to manage that risk with the executive team and have a plan in place."
Managing Editor Nicole Laskowski contributed to this report.
Disclosure: Informa holds a controlling interest in Informa TechTarget, which is the publisher of CIO Dive and the parent company of Omdia. Informa has no influence over CIO Dive's reporting.