Frankfurt am Main - November 2024

A Wildberries Architected Framework

In March, Iranian missiles turned two of the three Availability Zones in AWS’s Middle East (UAE) Region into a structural engineering problem. AWS’s health dashboard described it as “objects that struck the data center, creating sparks and fire." In reality, it was a targeted strike by Iran. The IRGC claimed responsibility citing AWS’s role hosting US military and intelligence networks, and Iranian state media went further, naming the US military’s use of AI systems running on AWS, Anthropic’s Claude included, for intelligence analysis and war simulations. Nearly six months later, ME-CENTRAL-1 and ME-SOUTH-1 are still, per AWS’s own service history, unable to reliably support customer applications.

Wearhouse on fire
Reported Ukrainian strike on a Wildberries
warehouse
(Exilenova Plus/Telegram)
In July, Ukrainian drones went after Wildberries, Russia’s logistics giant, essentially Amazon with worse labor conditions and better vodka. Twenty-plus warehouses hit since mid-July. Seven of the ten largest logistics centers disabled by mid-August. North of a billion dollars in direct losses to the company alone. Wildberries saw this coming and pre-negotiated the fallout. They had already rewritten their seller contracts to declare drone strikes force majeure.

All the while, Russia has been running its own operation against the rest of us. Drones over Polish, Romanian, German, and Baltic airspace are frequent enough that NATO shot down its fourth Romanian intruder of the year this month. Undersea cables severed by tankers that “accidentally” drag anchor for a hundred kilometers. Arson jobs farmed out over Telegram to randoms who don’t even know who’s paying them. IISS quantified the drone incursions this year: 144 suspected incursions across thirteen countries, all but one of them NATO members, and their polite term for how the alliance handled it was “a strategic failure." The stated goal, per the same report, is mapping air defense gaps and infrastructure weak points before anyone needs the map for real. Russia is engaging in reconnaissance by combat or target practice with plausible deniability.

Three events, one commonality: physical infrastructure is now a legitimate wartime target, and nobody pulling the trigger is checking whether your workload was Well-Architected first.

This Is On Your Side Of Shared Responsibility.

The Shared Responsibility Model for Resiliency dictates AWS owns the resiliency of the infrastructure, and each Region is built from multiple physically isolated Availability Zones specifically so a fault in one doesn’t propagate to the others. Fair, but what happens when multiple AZs are deliberately targeted? No amount of “Customer Obsession” will put me-central-1 together again to meet your RTO. Multi-AZ is a localized hardware failure mitigation. It was never a war plan.

The four DR strategies AWS actually gives you, backup and restore, pilot light, warm standby, multi-site active/active, sit on a cost curve for exactly this reason, and the framework’s own guidance is direct about when backup and restore stops being sufficient: the moment your definition of “disaster” moves from a bad terraform apply to Shahed drones targeting multiple AZs.

The Three Regions In The Cross Hairs

The three AWS regions closest to Russia are eu-north-1 (Stockholm), eu-central-1 (Frankfurt), and eusc-de-east-1, the new European Sovereign Cloud region in Brandenburg that went GA in January. If your workloads exist only in one of these, I’ve got bad news about what “closest to Russia” has come to mean in a year when Russia has been actively droning NATO airspace for sport.

Finding the data centers will be trivial. A facility pulling tens of megawatts continuously does not get to stay secret. It needs a dedicated grid interconnection, which means an application to Bundesnetzagentur or Svenska kraftnät that shows up in capacity planning paperwork somewhere. It needs cooling towers and generator farms visible on any halfway decent satellite pass, and trade press like Data Center Dynamics has been cataloging exactly this kind of hyperscaler campus across Europe for years. Unlike the US, it goes through local planning and environmental review with a public objection period. None of that requires cracking anyone’s beneficial ownership register, though for the record, Germany’s Handelsregister has been free and public since 2022. Sweden doesn’t bother with the half-measure. Its offentlighetsprincipen has kept government and property records unusually open since 1766.

A state intelligence service doesn’t need any of the above anyway. They’ve got SIGINT and HUMINT that skips every civilian step I just listed. The IISS report on the NATO drone campaign already told us the stated goal was mapping infrastructure weaknesses before anyone needs to act on the map. Assume the map already has your VPC on it.

Thundering Refugees Problem

Even if you do have a solid multi-region strategy, you’re not out of the woods. If any of the major AWS regions takes a drone strike, you and the rest of your availability zone neighbors will be clamoring for the same excess capacity in the other European regions. So it’s worth asking yourself which workloads must remain in the EU or EEA for regulatory reasons, and which workloads could move farther afield to the US, Canada, or South East Asia. If you do need to move within the EU, consider the logistics of regions like Spain or Milan. These may not yet be enabled for your accounts. These opt-in regions are less well known and so might have fewer refugees. They may also not have the same service availability as the regions you’re in.

Preparation here is key. You want to make sure that:

  1. You have a ranked list of regions you’ll try and migrate to
  2. The services you need are available in those regions
  3. Your account’s service quotas in those regions match those of your active region. The last thing you want in a recovery operation is to see “We’ve escalated your request to the service team”.

So What Do You Actually Do

While a multi-regional strategy seemed like overkill in 2019, if you’re operating in Europe and you don’t have a multi-regional distribution of your data and workloads, your business is at risk for getting drawn into this wider conflict.

A beautifully documented DR runbook that’s never been tested is worth exactly nothing the day mec1-az2 becomes a fire department problem instead of your TAM’s problem. If backup and restore genuinely covers what you run, fine, that’s a legitimate answer for your low-criticality workloads. But make sure it exists far enough away. If it doesn’t, warm standby or multi-site active/active isn’t the extravagance it was when eu-central-1 opened and the scariest thing on the horizon was a bad CloudFormation deploy. Stop treating “which Region” as a latency decision. It’s a geopolitics decision now, whether you asked for that or not.