On September 15, AWS wrote the sentence every cloud provider spends its whole existence trying to avoid: some customer data is gone, and there is no path back.
Two updates, posted the same day. In the UAE, AWS said it is "unable to restore access to the resources and data hosted exclusively in mec1-az2," one of three availability zones in its me-central-1 region; work continues on the other two, mec1-az1 and mec1-az3, and on the shared regional infrastructure around them. In Bahrain, the loss is total. AWS said it is "unable to restore access to the resources and data hosted exclusively in this Region," after damage that spread across multiple availability zones and, in the company's own words, "exceeded what our regional and multi-AZ services are designed to withstand."
This is the same story 19 Days Dark touched in passing: strikes during the Iran conflict hitting AWS facilities in the UAE and Bahrain, a 48-hour outage that took two Abu Dhabi banks offline in March. Two AWS facilities in the UAE were struck directly. In Bahrain, a drone strike close to one facility caused structural damage, disrupted power delivery, and forced fire suppression that added water damage on top of it. What that piece didn't yet know, because nobody did, is that a second Bahrain availability zone went down in April, taking the entire region offline, and that AWS would spend the following five months trying to recover it before saying, in September, that it had stopped trying. Not down anymore. Gone.
→ The enterprise continuity story this event started, in full: 19 Days Dark.
In June, the question was how fast a vendor could bring something back. By September, for two facilities in the Gulf, the honest answer became: it can't.
What multi-AZ actually promises
AWS's own explanation is the most useful sentence in the whole disclosure: the damage "exceeded what our regional and multi-AZ services are designed to withstand." Multi-availability-zone redundancy is built for the failure mode cloud engineers spend their careers guarding against: a bad deployment, a power failure, a fire in one building. It assumes that whatever takes out zone A leaves zone B and zone C alone, because they sit in physically separate facilities. That assumption held for two of the three UAE zones, still in active recovery five months on. It did not hold for the third, or for the whole of Bahrain, because the actual failure mode wasn't a single-facility accident. It was a regional armed conflict, capable of reaching more than one facility in the same footprint within the same season, regardless of which zone label was printed on the rack.
Most customers in both regions had already moved on before the loss became permanent, AWS said, "using backups where available or implementing alternative solutions." That line does as much work as the outage itself. It means the difference between a bank that lost nothing and a bank that lost something permanently wasn't whether it had a backup. It was whether that backup's risk profile actually differed from the primary's, or just carried a different zone label inside the same blast radius.
The line DORA already drew
European banks have had a name for this distinction since January 2025; they just hadn't had a live example to point to. The EU's Digital Operational Resilience Act requires financial entities to test backup and recovery plans against switchover scenarios at least yearly, and its Article 12 goes further for one category of institution: a central securities depository's secondary processing site must sit "at a geographical distance from the primary processing site to ensure that it bears a distinct risk profile." That specific wording is scoped to CSDs, not every bank. But the logic behind it isn't. A second copy in a second availability zone of the same region does not automatically have a distinct risk profile from the first. Whether it does is an empirical question, and Iran's strikes just answered it for two of AWS's three UAE zones, and for the whole of Bahrain.
→ The same conflict, covered from the sovereignty angle, chip access as a foreign-policy reward: The Rented Border.
What a contract couldn't have covered anyway
DORA's Article 29 asks banks to weigh concentration risk before they sign: is this provider, or this region, "easily substitutable." Any institution that ran a critical function exclusively out of mec1-az2 or out of Bahrain, with no failover outside that footprint, had already answered that question without pricing in what the answer meant. Article 30 then governs what the contract says afterward, a guaranteed exit transition, a provider's commitment to return customer data on termination. Every affected customer presumably had language like that. None of it mattered here, because an exit clause assumes the data still exists somewhere to hand back. It's written around a provider walking away, not a provider's facility taking a direct hit. A perfectly compliant exit clause is worth nothing if there's no copy left to return, which is exactly why the real defense was never going to live in the contract.
A backup that shares its primary's blast radius was never really a second copy. It was the same bet, filed twice.
What holds up
None of this is a job Appice's Traffic Manager was built for; that pattern solves reversibility at a different layer, which delivery channel a message travels through when a provider degrades, not where a bank's core data physically sits. But the discipline underneath is the same one this event just tested for real: a fallback only counts if whatever could plausibly damage the primary genuinely can't reach it too. For the channel layer, that's a provider monitored continuously and substituted in under an hour, the pattern covered in 19 Days Dark. For infrastructure and data, it's a sharper version of the question most vendor reviews still stop short of: not just where is the backup, but how many kilometers away, how many facilities apart, and could one event, a war, a grid failure, a fire, plausibly reach both at once.
One more question worth adding to a vendor review after this specific incident: when a provider says a region has three availability zones, ask how far apart they actually sit, and what kind of event, short of a meteor, could plausibly take out more than one. Bahrain's answer turned out to be a single regional conflict, close enough to reach every zone it had.
AWS didn't break a promise in Bahrain. It kept the only one it could: rebuild what can be rebuilt, and say plainly when something can't. The promise a bank actually needed, that no single event could reach every copy of its data at once, was never the vendor's to make. It was the bank's to build.
The full argument for reversible architecture, the Traffic Manager pattern included, is in The Perimeter.