Incident Response Flowchart
The code
flowchart TD
A[Alert fires] --> B{Severity high?}
B -- No --> C[Create ticket]
B -- Yes --> D[Page on-call engineer]
D --> E[Acknowledge within 5 min]
E --> F{Customer impact?}
F -- Yes --> G[Post status update]
F -- No --> H[Investigate quietly]
G --> I[Identify root cause]
H --> I
I --> J{Fix verified?}
J -- No --> I
J -- Yes --> K[Write postmortem]
K --> L((Incident closed))
How this template works
This template is the flowchart equivalent of an on-call runbook. An alert arrives, a human is paged or a ticket is filed, the responder acknowledges, communicates, investigates, and finally closes the incident with a postmortem. Teams paste it into their incident tooling docs, print it for the on-call binder, and walk new engineers through it during handover. The branches encode policy decisions that are easy to forget at three in the morning: which severities page a human, when to post a public status update, and what happens when a fix does not actually work.
Reading the code top to bottom: flowchart TD gives a top-down layout that mirrors the timeline of a real incident. A[Alert fires] is the entry rectangle, and B{Severity high?} is the first policy gate — the No branch files a quiet ticket while the Yes branch pages the on-call engineer. The step E[Acknowledge within 5 min] encodes an expectation, not just an action, which is a useful trick: putting the target directly in the label means nobody has to remember it. The diamond F{Customer impact?} splits the communication duty from the investigation duty, and both paths converge on I[Identify root cause]. The most important line is the back edge J -- No --> I: if the fix is not verified, investigation resumes. Without that loop the diagram would quietly claim that every first attempt succeeds.
The gotcha here is the loop target. A back edge must point at a node id that already exists — write J -- No --> I and the parser connects to the investigation node defined earlier; misspell the id and Mermaid silently creates a brand-new empty node, which shows up as a mystery box floating below the flow. Also watch label punctuation: E[Acknowledge within 5 min] is safe, but E[Acknowledge within 5 min (or escalate)] needs quotes around the whole label because of the parentheses.
To adapt it, replace the severity question with your actual paging policy, and add a node for your incident channel or war room if you use one. If your team has a separate communications role, wrap the status update step in a subgraph to make ownership obvious.
Related templates: the fault handling flowchart models the technical retry logic that happens inside the investigation step, the decision tree flowchart shows how to route alerts before they ever page anyone, and the CI/CD pipeline flowchart covers the deploy events that often trigger incidents. The flowchart diagram guide has the full syntax reference.
Variations to try
- Add a severity diamond for medium alerts that routes to a business-hours rotation instead of paging.
- Insert a comms subgraph around the status update step for teams with a dedicated communications role.
- Duplicate the diagram per service and change the alert label to the actual monitoring source.