00:00
System — Incident opened. Scenario: AWS eu-west-1 region failure. Services affected: checkout, auth, media-cdn.
00:01
Sarah — Bridge is now open. All teams please confirm attendance. Severity is P1. We are running ITIL Major Incident protocol.
00:02
Emma — Paging SRE Lead, Problem Manager, and Change Manager. Stakeholder notification going out now to C-suite.
00:04
Niamh — Problem Manager on bridge. Logs show eu-west-1 AZ degradation starting 00:00 UTC. Likely EC2 capacity event. Opening Problem Record PR-0041.
00:07
You — Approved failover to eu-central-1. SRE team to proceed with Route 53 weight shift.action taken
00:09
Conor — EBF-0091 raised for Route 53 weight shift. Awaiting IM approval before proceeding.
00:11
You — EBF-0091 approved.approved
00:13
Aoife — Route 53 weight shift to eu-central-1 complete. Monitoring error rates. Checkout latency still elevated at 4.2s p99.
00:18
System — SLA breach: Time-to-workaround exceeded 15-min threshold.SLA breach
00:21
Emma — Second stakeholder update sent. Estimated resolution window communicated as 30 mins.
00:23
Aoife — Error rates dropping. p99 now 1.8s. eu-central-1 handling full traffic load. Recommend holding 5 mins before declaring workaround in place.