

0 / 2 embers
0 / 3000 xp
click for more info
Complete a lesson to start your streak
click for more info
Still calibrating
click for more info
Not enough gems
Cost: 6 gems
1: Observability
incomplete
2: Alerts
incomplete
3: Responsible Disclosure
incomplete
4: Incident Severity and Triage
incomplete
5: Damage Control
incomplete
6: Postmortems
incomplete
7: Incident Reporting
incomplete
Back
ctrl+,
Next
ctrl+.
This lesson's interactive features are locked, please to keep using them
Once you realize you have a problem, ideally through an alert (less ideally through a news headline), you need to triage it.
You gotta figure out how bad it is quickly enough to respond proportionally. Not every anomaly is an incident, and not every incident is "critical."
In movies, you've probably heard people shout "It's a Code Red!" or "We're escalating to DEFCON 3!" Those are just labels for "severity levels".
Most organizations use 3–5 severity levels. The exact labels don't matter too much. Some companies use colors (yellow, orange, red), some use severity numbers (sev3, sev2, sev1), and some use descriptive labels (medium, high, critical). The important thing is that everyone understands what each level means and how to respond.
When something looks wrong, ask three questions:
Those answers determine severity.
It's best practice to assign a single person (dramatically dubbed the incident commander by people who have too much time on their hands) when an incident is discovered. That person doesn't need to be the hero who fixes the problem, but they do need to coordinate the response, keep communication flowing, and make sure it gets resolved.
Without an owner, multiple people investigate the same thing, nobody updates stakeholders, and fixes get delayed because everyone assumes someone else is handling it.