Establish practice
Define cadence, owner, participation, and success signals for a recurring team practice.
Establish the practice
Define cadence, owner, participation, and success signals so the method works as a recurring team practice.
Variant for Operating practice. The plan is derived from method form, method steps, and runsheet.
RunsheetAn operating practice works through repetition. The cadence should be small enough to hold reliably.
The steps describe one repetition and are adjusted after review.
- 1
Phase 1: Activation
0-5 minFirst responder takes IC role and posts in the channel: I am IC, severity X, hypothesis Y. Update channel topic with IC name. Confirm pager alert.
OwnerIf no one takes IC, pager escalates after 5 min. IC take over must be explicit, not implicit. No "someone will handle it". - 2
Phase 2: Assign roles
5-15 minIC calls Operations Lead and Communications Lead. Scribe is named. SMEs are paged selectively, not everyone. Start conference bridge and keep channel as primary communication medium.
TeamIC stays coordinator, not operator. If IC spends all time clicking, no one holds the overall view. For SEV1, assign OL immediately. - 3
Phase 3: Investigation and mitigation
variableOL runs hypothesis tests, scribe records timestamps. IC coordinates: what has been tested, next action, owner. CL posts status updates every 30 min to stakeholders.
TeamMitigation takes priority over root-cause identification. If rollback is available, OL runs rollback and causes analysis follows in postmortem. No parallel actions without IC confirmation. - 4
Phase 4: Resolution and handoff
15 minIC confirms recovery based on SLO metrics, not gut feeling. Move channel topic to Resolved. Close status-page entry. For long-running incidents, perform handoff to new IC every 2 h.
TeamRecovery needs two consecutive green measurement intervals. A single green tick is not enough. Handoff requires a written statement (status, open points, next actions). - 5
Phase 5: Transition to postmortem
15 min after resolutionIC creates postmortem ticket with incident ID, links channel log and timeline. Schedule postmortem within 5 working days. Collect action items in ticket before closure.
TeamAnyone leaving channel before postmortem ticket exists risks losing lessons. Scribe is responsible for complete log export.
Practice Charter
Copyable charter with purpose, cadence, owner, participation, flow, success signals, and review.
practice-charter.md
Practice Charter: Incident Command
Purpose
The method helps clarify incident work, roles, and countermeasures concretely. It aligns incident picture, responsibilities, and next countermeasures. The outcome is captured as incident log, action tracker, and stakeholder updates.
Cadence
weekly
Owner and Participation
- Owner: open
- Participants: 4-15
- Mode: synchronous
Flow per Iteration
- Phase 1: Activation (0-5 min) First responder takes IC role and posts in the channel: I am IC, severity X, hypothesis Y. Update channel topic with IC name. Confirm pager alert.
- Phase 2: Assign roles (5-15 min) IC calls Operations Lead and Communications Lead. Scribe is named. SMEs are paged selectively, not everyone. Start conference bridge and keep channel as primary communication medium.
- Phase 3: Investigation and mitigation (variable) OL runs hypothesis tests, scribe records timestamps. IC coordinates: what has been tested, next action, owner. CL posts status updates every 30 min to stakeholders.
- Phase 4: Resolution and handoff (15 min) IC confirms recovery based on SLO metrics, not gut feeling. Move channel topic to Resolved. Close status-page entry. For long-running incidents, perform handoff to new IC every 2 h.
- Phase 5: Transition to postmortem (15 min after resolution) IC creates postmortem ticket with incident ID, links channel log and timeline. Schedule postmortem within 5 working days. Collect action items in ticket before closure.
Success Signals
Incident Log, Action Tracker, Stakeholder Updates
Review
4 weeks