A simulated incident, announced in advance, run against production with a deliberately broken dependency.
the drill: the supplier API returns 503 for 20 minutes,
simulated by a firewall rule on the egress host.
what the responder had to do:
notice it (the alert fired in 4 minutes — good)
identify it (8 minutes — the runbook's first step
pointed at the wrong dashboard)
enable the degraded path (2 minutes)
communicate (not done — the status page step is
still being skipped)
total: 14 minutes to mitigation, and one runbook fix.
Announcing it in advance removes the value of surprise and removes the risk of a drill becoming a real incident, which is the right trade when the goal is exercising a procedure rather than testing vigilance. The skipped communication step is the finding that keeps recurring and is now the first step rather than the ninth.