Everything an on-call engineer needs the moment something is on fire: how to classify it, who owns what, the six phases from detection to postmortem, and the exact messages to send. Read it now, print it for the wiki, fork it for your stack.
Runs the incident, makes the calls, keeps the timeline. Does not fix things personally — coordinates those who do.
Drives technical investigation and remediation. Reports status and blockers to the IC, not the channel.
Drafts internal and external updates from templates. The single voice to customers, status page, and support.
Timestamps every decision, action, and finding in the incident channel. The raw material for the postmortem.
🚨 [SEV-X] declared — [one-line summary] IC: @name · Ops: @name · Comms: @name Impact: [systems / customers affected] Status: investigating. Updates every [15 / 30] min in this thread.
We are investigating reports of [degraded performance / errors] affecting [product area]. Some users may experience [symptom]. Our team is engaged and the next update will be posted by [time + TZ].
Hi [name], We're writing to let you know about an incident that may have affected your account between [start] and [end] ([TZ]). What happened: [plain-language description] What was affected: [data / service] What we've done: [containment + fix] What you should do: [action, or "no action needed"] We take this seriously and will follow up with our full review.
This incident has been resolved as of [time + TZ]. The issue was [root cause in one sentence]. All services are operating normally. A full postmortem will follow within [5 business days].
Uproot continuously collects the access logs, change records, and alerts your auditor asks for after an incident — so when CC7.4 comes up, you're pulling a link, not reconstructing a timeline from Slack.