methodatlas
All topicsTopic

reliability

3 methods with this tag.

Related topics
DevOps

Blameless Postmortem

Turns incident work, roles, and countermeasures into a tangible result by documenting the incident, describing impact and timeline, and sharing learnings.

Core question: Which system conditions and decision points made the incident possible, and which concrete measures prevent recurrence?

Operating practiceMediumWorkshop + async
30-90 min
3-12

Run sheet · visual · session plan

DevOps

Incident Command

Turns incident work, roles, and countermeasures into a tangible result by declaring an incident, assigning commander and roles, and closing and reviewing the incident.

Core question: Who is in each role, which effect should be checked next, and when is the next stakeholder update due?

Operating practiceMediumWorkshop + async
As needed
4-15

Run sheet · visual · session plan

Operations

Runbook

Turns workflows, data, causes, and improvements into a tangible result by defining a scenario, writing steps, and testing and updating.

Core question: Which steps does a trained person execute in which order to handle the trigger safely without having to improvise?

Operating practiceLowAsync
20-60 min
1-4

Run sheet · visual · session plan