UprootSecurity
Book a demo
RunbookSOC 2 · CC7.4Editable

Incident response runbook

Everything an on-call engineer needs the moment something is on fire: how to classify it, who owns what, the six phases from detection to postmortem, and the exact messages to send. Read it now, print it for the wiki, fork it for your stack.

Format
Runbook · printable
Sections
Severity · Roles · 6 phases · Comms
Downloads
2,810
Section 01

Purpose & scope

This runbook governs how [YOUR COMPANY] detects, responds to, and learns from security and availability incidents. It exists so that nobody has to improvise judgment at 3 a.m.
It applies to any event that compromises — or credibly threatens — the confidentiality, integrity, or availability of production systems, customer data, or the corporate environment. When in doubt, declare an incident. Over-declaring is cheap; under-declaring is how a SEV-3 becomes a breach notification.
Activation: anyone may declare an incident. Page the on-call via [PAGERDUTY / OPSGENIE], post in #incidents, and the responder becomes Incident Commander until explicitly handed off.
Section 02

Severity matrix

Classify within the first five minutes. Severity sets the response SLA, who gets paged, and whether the clock starts on regulatory notification.
SEV-1
Critical
Confirmed data breach, full outage, or active intrusion. Customer data or trust at stake.
Page immediately · IC + exec + legal · Ack < 5 min
SEV-2
High
Major degradation, suspected breach, or a control failure with real exposure. No confirmed data loss yet.
Page on-call · IC + eng lead · Ack < 15 min
SEV-3
Moderate
Contained issue, single-tenant impact, or a vulnerability needing urgent patching. Workaround exists.
Business hours · Owning team · Ack < 1 hr
SEV-4
Low
Minor anomaly, near-miss, or informational alert worth tracking. No active impact.
Next business day · Ticket & triage
Section 03

Roles & on-call

One person can wear several hats on a small incident, but the roles never disappear — someone owns each. The Incident Commander assigns them explicitly and out loud.
Owns the response

Incident Commander

Runs the incident, makes the calls, keeps the timeline. Does not fix things personally — coordinates those who do.

Owns the fix

Operations Lead

Drives technical investigation and remediation. Reports status and blockers to the IC, not the channel.

Owns the message

Communications Lead

Drafts internal and external updates from templates. The single voice to customers, status page, and support.

Owns the record

Scribe

Timestamps every decision, action, and finding in the incident channel. The raw material for the postmortem.

Section 04

The six phases

Detection through closure. The phases overlap in practice, but the order of priorities does not: stop the bleeding before you explain it.
1

Detect & declare

Goal: a named incident with a severity and an IC
  • Confirm the signal is real, not a flapping alert
  • Assign severity from the matrix above
  • Open the channel, page responders, name the IC
2

Triage & assess

Goal: known blast radius
  • What systems, data, and tenants are affected?
  • Is it ongoing or contained? Is data leaving?
  • Start the timeline — when did it actually begin?
3

Contain

Goal: stop the spread
  • Isolate affected hosts, revoke compromised credentials
  • Preserve evidence before you wipe anything
  • Apply the least-destructive control that works
4

Eradicate & recover

Goal: clean systems back in service
  • Remove the root cause, not just the symptom
  • Rebuild from known-good, rotate all related secrets
  • Verify the fix in production before declaring recovery
5

Notify

Goal: the right people told, on time
  • Check contractual and regulatory notification clocks (GDPR 72h, customer SLAs)
  • Loop in legal before any external statement
  • Update the status page and affected customers
6

Review

Goal: it cannot happen the same way twice
  • Blameless postmortem within 5 business days
  • Every action item has an owner and a date
  • Feed findings back into detection and this runbook
Section 05

Comms templates

Fill in the brackets and send. Pre-written so the Communications Lead is editing, not composing, while the clock runs.
Internal — incident declared
Post in #incidents at declaration
🚨 [SEV-X] declared — [one-line summary]
IC: @name  ·  Ops: @name  ·  Comms: @name
Impact: [systems / customers affected]
Status: investigating. Updates every [15 / 30] min in this thread.
External — status page (investigating)
Within SLA of declaration
We are investigating reports of [degraded performance / errors] affecting [product area]. Some users may experience [symptom]. Our team is engaged and the next update will be posted by [time + TZ].
Customer — confirmed impact
After legal review, SEV-1/2
Hi [name],

We're writing to let you know about an incident that may have affected your account between [start] and [end] ([TZ]).

What happened: [plain-language description]
What was affected: [data / service]
What we've done: [containment + fix]
What you should do: [action, or "no action needed"]

We take this seriously and will follow up with our full review.
External — resolved
After recovery verified
This incident has been resolved as of [time + TZ]. The issue was [root cause in one sentence]. All services are operating normally. A full postmortem will follow within [5 business days].
Section 06

Postmortem outline

Blameless, factual, and shipped within five business days while memory is fresh. Copy this structure into a doc and fill it from the scribe's timeline.
1

Summary

2–3 sentences a non-engineer can read
2

Impact

Who/what was affected, for how long, measured
3

Timeline

Detection → declaration → containment → recovery, timestamped
4

Root cause

The real why — five whys, not the proximate trigger
5

What went well / what didn't

Honest, blameless, specific
6

Action items

Each with an owner, a due date, and a tracking link
The test of a good postmortem:every action item closes a gap that, if it had existed, would have prevented or shortened this incident. If an item wouldn't have helped, cut it.
Section 07

First-15-minutes checklist

When you're the one who gets paged. Print this part and tape it to the wall.
  • Acknowledge the page so others know it’s owned
  • Confirm the signal is real — reproduce or verify it
  • Assign a severity from the matrix
  • Open #incidents and declare; you are IC until handoff
  • Name Ops, Comms, and Scribe (even if it’s just you for now)
  • Establish blast radius: systems, data, tenants
  • Decide: contain now, or investigate first?
  • Start the timeline — note when it began, not just when you noticed
  • If data may have left: flag legal and the notification clock
  • Post the first status update; set the next-update time

A runbook is only as good as the evidence behind it

Uproot continuously collects the access logs, change records, and alerts your auditor asks for after an incident — so when CC7.4 comes up, you're pulling a link, not reconstructing a timeline from Slack.

Book a 20-min demoGet the readiness checklist