loading…
Detect → Triage → Mitigate → Recover → Understand → Prevent
Mitigation reduces current harm. Root-cause analysis explains the failure. A permanent fix changes the system so recurrence is less likely. Do not delay mitigation while searching for a perfect explanation.
Confirm customer impact, declare severity, assign an incident lead, stop risky changes, choose a reversible mitigation, establish a communication cadence, and preserve evidence. One person coordinates; responders should not all investigate the same clue.
Choose either an SDK regression or Kintsu issuing duplicate refunds to 3% of requests. Write the first status update, immediate containment, owner assignments, evidence to preserve, and recovery criteria.
State what users experience, what is known, what is being done, and when the next update will arrive. Do not fill uncertainty with guesses.