Beyond the 30-Minute AZ Test: Running Multi-Day AWS Availability Zone Evacuation Drills

Beyond the 30-Minute AZ Test: Running Multi-Day AWS Availability Zone Evacuation Drills

Cloud Platform Engineering

Short AZ evacuation tests catch the obvious failures — applications that haven’t configured multi-AZ deployment, load balancers that don’t drain correctly, health checks that time out. Fifteen minutes of shifted traffic proves your routing rules work and surfaces hard failures.

Multi-day evacuations prove something harder: that your infrastructure remains stable when the evacuated state becomes the permanent state. The failures that emerge after 24–72 hours are different in kind from the ones that appear in the first hour, and they’re the ones that matter most for actual availability.

Why Duration Changes What You Find

Auto Scaling Capacity Drift

When an AZ is evacuated, Auto Scaling groups redistribute capacity to remaining zones. A 15-minute test verifies that this redistribution happens. A 72-hour test verifies that the redistributed capacity holds.

The failure mode: Auto Scaling launches instances in healthy AZs to compensate, but the group’s DesiredCapacity settings, lifecycle hooks, and cooldown periods may be tuned for normal balanced distribution. Under sustained single-AZ evacuation, scaling events that fire during the day — scheduled policies, predictive scaling adjustments, SQS-triggered scaling — can behave unexpectedly when the AZ balance assumption no longer holds.

Key metric to watch: GroupDesiredCapacity vs GroupInServiceInstances per AZ. Divergence indicates capacity that was expected but not provisioned.

DNS TTL Violations

Route 53 health checks and DNS failover work during the initial AZ evacuation. The TTL violation that emerges over hours: clients that cache DNS records beyond their published TTL, sending traffic to the evacuated endpoint after the DNS has updated.

This is an application-layer issue, not an AWS issue. Client libraries, connection pools, and service mesh sidecar configurations all have their own interpretation of DNS TTLs. Multi-day tests reveal which clients aren’t respecting the cache headers they receive.

Connection Pool Exhaustion

Connection pools are sized for expected traffic patterns. When an AZ is evacuated and its traffic redistributes to remaining instances, the per-instance connection pool fills with traffic that was previously distributed across more endpoints. Pools sized for normal-distribution traffic are undersized for single-AZ concentration.

The specific failure pattern: connection pools that report healthy during the first few hours, as requests complete and connections release normally, but begin showing timeout errors after 12–24 hours as the queue depth slowly climbs. Connection pool metrics — ActiveConnections, DatabaseConnections, ConnectionPoolAvailableConnections — need longer observation windows to show this drift.

Deployment Pipeline Failures

Routine deployments fire on their normal schedule regardless of evacuation state. A deployment pipeline that was written without awareness of the evacuation may behave unexpectedly: attempting to deploy to capacity in the evacuated AZ, triggering rollback logic that itself fails on the skewed AZ distribution, or updating configuration parameters that override the evacuation routing.

This failure mode only appears if you run a real deployment during the evacuation window — something a 30-minute test won’t catch.

Component-Specific Runbooks

ECS with Application Load Balancer

Before evacuation:

aws arc initiate-zonal-shift \
  --resource-identifier {ALB_ARN} \
  --away-from az-us-east-1a \
  --expires-in 72h \
  --comment "multi-day drill 2026-10-02"

Monitor during evacuation:

  • TargetResponseTime per AZ on the ALB
  • ECS ServiceDeploymentCircuitBreaker state (should be DISABLED during drill)
  • ECS DesiredCount vs RunningCount per task definition
  • Container instance health per remaining AZ

Key validation: Deploy one service change during the 72-hour window. Verify the deployment completes without touching the evacuated AZ and that rollback behavior works correctly under the shifted traffic distribution.

EKS with Network Load Balancer

EKS zonal shift requires the AWS Load Balancer Controller and topology-aware routing configured on the services:

# Service annotation for NLB zonal shift support
annotations:
  service.beta.kubernetes.io/aws-load-balancer-type: "external"
  service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: "ip"
  service.beta.kubernetes.io/aws-load-balancer-cross-zone-load-balancing-enabled: "false"

Cross-zone load balancing disabled is the correct configuration for zonal shift — it ensures the NLB only routes to healthy AZs rather than distributing across all registered targets regardless of zone.

Monitor during evacuation:

  • Node count per AZ in the evacuated state
  • Pod scheduling behavior — kube-scheduler should place pods on nodes in healthy AZs only
  • kube_pod_info metrics by node zone label
  • HPA scale events: verify Horizontal Pod Autoscaler fires correctly under zone-concentrated load

Known failure mode: Node groups sized with a minimum count per AZ may attempt to maintain capacity in the evacuated AZ, fighting the zonal shift. Review minSize in the node group configuration before starting the drill.

RDS Multi-AZ

RDS Multi-AZ with ARC Zonal Shift evacuates the standby replica when the primary is in the healthy AZ, or promotes a failover when the primary’s AZ is evacuated.

For a drill targeting the primary AZ:

aws rds failover-db-cluster \
  --db-cluster-identifier {cluster-id}
# Then initiate zonal shift on the original primary AZ

Monitor during evacuation:

  • Replication lag on the new standby (original primary zone)
  • Connection counts to primary vs replica endpoints
  • FreeStorageSpace on both instances (replication backlog under load can consume storage)
  • Maintenance window behavior: RDS automatic maintenance will attempt to apply patches during the maintenance window regardless of drill state. Confirm or postpone the next maintenance window before starting a multi-day drill.

Aurora PostgreSQL

Aurora’s multi-AZ behavior with zonal shift is more complex because Aurora’s read replica fleet participates in zonal routing:

With Aurora Global Database: The zonal shift applies to the regional cluster. Cross-region reader endpoints are unaffected by intra-region zonal shifts — verify that application read routing correctly distributes between regional readers in healthy AZs and avoids the evacuated AZ’s reader endpoint.

Writer endpoint during evacuation: Aurora’s writer endpoint routes to the primary instance. If the primary is in the evacuated AZ, initiate a manual failover to promote a reader in a healthy AZ before the zonal shift:

aws rds failover-db-cluster \
  --db-cluster-identifier {aurora-cluster-id} \
  --target-db-instance-identifier {reader-in-healthy-az}

Monitor during evacuation:

  • AuroraReplicaLag per replica (should remain <100ms under normal load)
  • CommitLatency on the writer (sustained single-AZ evacuation shouldn’t affect writer latency, but connection pool concentration from evacuated-AZ application instances can)
  • NetworkReceiveThroughput per AZ — asymmetric cross-AZ traffic has cost implications for multi-day drills

What a Multi-Day Drill Proves

A 72-hour successful evacuation drill produces three categories of evidence:

Capacity sufficiency: The remaining AZs can sustain full production traffic load for a duration that exceeds most real-world AZ impairments. AWS AZ incidents lasting more than 12–24 hours are extremely rare — a 72-hour drill proves you can handle any realistic event.

Database stability: Connection pool sizing, replication health, and maintenance window behavior are validated under sustained single-AZ load. This is the category where most surprises appear.

Operational readiness: Your deployment pipeline, monitoring alerts, on-call runbooks, and capacity tooling all behave correctly during an active evacuation. The drill is also a validation of your operational documentation.

The So What

Multi-day AZ evacuation drills are distinct from chaos engineering experiments — they’re closer to fire drills that produce auditable evidence. Compliance frameworks increasingly require documented proof of availability architecture, not just architectural diagrams. A successful 72-hour drill with metrics and deployment events is the highest-quality evidence available.

For teams that have run short AZ tests but haven’t sustained them: the failure modes above — capacity drift, TTL violations, connection pool exhaustion, maintenance window collisions — are all real incidents from production AZ impairments. The 30-minute test doesn’t find them. The multi-day drill does.

Content created with AI assistance and reviewed for accuracy.

💬

Join the conversation

Stack Insiders is our free community for readers who want to go deeper — share resources, ask questions, and connect with others across every vertical we cover.

Join Stack Insiders →

Newsletter coming soon.

Curated digests across AI, biohacking, photography, travel, and more. Be the first to know when we launch.