methodatlas
Playbook

Systematically Improve Engineering Quality

Move from a vague quality goal through visible flow and tested hypotheses to a durably anchored standard path.

Steps5 methods
Time2-4 weeks
FormatHybrid
Outcome

A confirmed cause of recurring quality problems and a documented, proven standard path to avoid them.

At the end you have

GQM planKanban board with flow metricsTested hypothesesIsolated fault sourceGolden path document

Decision point

You can decide which standard path becomes mandatory and how it is anchored in the team.

Next step

Re-check the metrics from the GQM plan after four weeks to confirm the effect of the golden path.

Ideal for

  • Recurring quality problems in engineering
  • Unclear causes behind poor quality metrics
  • Teams without an established standard path for common technical tasks

Not good for

  • One-off, clearly localized bugs
  • Quality problems without any available metric or observation
Preparation

What should be clear before you start

Roles

  • Tech lead or engineering manager
  • Involved development team

Inputs

  • Existing quality metrics or incident history
  • Access to the current Kanban board

Setup

  • Gather available metrics upfront
  • Plan a realistic time frame for the cause investigation
Flow

Method path

5 methods
  1. 1EngineeringGQM plan

    Goal Question Metric

    Which questions must be answered to assess the defined goal from the chosen viewpoint, and which metrics provide those answers?

    Why this step?

    GQM prevents quality from being treated as a vague feeling and makes progress measurable.

    The defined metrics then show where the workflow actually stalls.

    Paper illustration of Goal Question Metric with its method-specific working model.
  2. 2EngineeringKanban board with flow metrics

    Kanban

    What helps us finish started work promptly and more predictably within our agreed workflow?

    Why this step?

    Visible flow shows at which point in the process quality problems systematically arise or pile up.

    Conspicuous flow patterns provide the basis for concrete, testable cause hypotheses.

    Paper illustration of a four-column Kanban board with limited ongoing work, a visible blocker and a review loop.
  3. 3EngineeringTested hypotheses

    Hypothesis-Driven Troubleshooting

    Which competing hypothesis best explains all observed signals, and which safe test discriminates among them most strongly?

    Why this step?

    The hypothesis-driven approach prevents premature fixes and ensures the real cause gets tested.

    A confirmed hypothesis is narrowed down technically to the exact fault source.

    Paper illustration for Hypothesis-Driven Troubleshooting.
  4. 4EngineeringIsolated fault source

    Fault Isolation

    Which smallest subsystem contains the fault when candidates are systematically excluded through discriminating tests?

    Why this step?

    Fault isolation ensures the fix addresses the actual technical source, not a symptom.

    From the resolved cause a proven standard path for future work can be derived.

    Paper illustration for Fault Isolation.
  5. 5EngineeringGolden path document

    Golden Path

    How does a team reliably move from entry to an operable result in the supported standard case?

    Why this step?

    A documented golden path prevents the same quality problem from recurring elsewhere.

    The golden path is anchored in the team and regularly checked against new metric values.

    Paper illustration for Golden Path
Completion criteria
Templates

Artifacts for this playbook

Artifacts stay collapsed until you actually need them.

MarkdownShow template

Goal Question Metric: Worksheet

Plan Goal Question Metric around its domain steps and the concrete decision question.

# Goal Question Metric: Worksheet

## Question
To be confirmed

## Desired outcome
To be confirmed

## Scope
Full workshop
Time: 60–90 minutes

Complete this worksheet on paper or in your own document during the work.

## Preparation for this scope

### Method setup
- **Materials:** Measurement object and purpose, Quality focus and viewpoint, Context and available data, shared workspace, source links, and decision log.
- **Roles:** Goal owner · domain expert · measurement/analytics · data owner
- **Advance information:** Collect Measurement object and purpose, Quality focus and viewpoint, Context and available data in advance and mark open assumptions.
- **Overall time needed:** 60–90 minutes
- **Setup:** Prepare Goal → Questions → Metrics with an explicit interpretation model visibly and calibrate assessment terms before starting.

### Prepare for the work steps

#### GQM goal: GQM goal
- Measurement object and purpose
- Quality focus and viewpoint
- Context and available data

#### Questions: Questions
- GQM goal
- Quality focus and viewpoint
- Context and available data

#### Metrics: Metrics
- question tree
- Quality focus and viewpoint
- Context and available data

#### Interpretation: Interpretation
- metric set
- Quality focus and viewpoint
- Context and available data


## GQM plan

### Goal
Object · purpose · quality focus · viewpoint · context

...

### Questions
Which answers assess the goal?

...

### Metrics
Formula · unit · source · segment · quality

...

### Interpretation model
How do values yield a domain decision?

...

Enter only real input, sources, and decisions.

[NASA Software Engineering Handbook: The Goal Question Metric Approach](https://swehb.nasa.gov/spaces/7150/pages/16450343/SWEREF-391)


## GQM goal: GQM goal
Expected artifact: GQM goal

GQM goal · source · uncertainty · next decision

Entry:

...

- [ ] Are all five goal elements unambiguous?

## Questions: Questions
Expected artifact: question tree

question tree · source · uncertainty · next decision

Entry:

...

- [ ] Does every question cover a relevant aspect of the goal?

## Metrics: Metrics
Expected artifact: metric set

metric set · source · uncertainty · next decision

Entry:

...

- [ ] Can every metric answer at least one question?

## Interpretation: Interpretation
Expected artifact: interpretation model

interpretation model · source · uncertainty · next decision

Entry:

...

- [ ] Is the combination of metrics into a decision traceable?

## Open questions and next steps

...

Method guide: https://methodatlas.meierhoff-systems.de/en/methods/goal-question-metric/run-sheet
MarkdownShow template

Kanban: Worksheet

Plan setup, initial trials and regular reviews. Select one step to prepare a specific appointment.

# Kanban: Worksheet

## Question
To be confirmed

## Desired outcome
To be confirmed

## Scope
Set up and continue Kanban
Time: Several setup steps, followed by ongoing management and improvement

Complete this worksheet on paper or in your own document during the work.

## Preparation for this scope

### Method setup
- **Materials:** Shared board, representative work items, visible policies and start and finish dates. A simple physical board can support the initial work.
- **Roles:** People in the workflow · ownership of policies and improvement · stakeholders as needed
- **Advance information:** Gather real items, actual waiting and handover steps, known blockers and available time data.
- **Overall time needed:** Setup and initial trials across several appointments; ongoing practice afterward.
- **Setup:** Define work item types, start and finish points, relevant states, WIP control and transition policies. Establish a Service Level Expectation (SLE) with elapsed time and probability, initially a clearly labeled estimate and later based on historical Cycle Times. Limits fit your workflow and are reviewed; no fixed team-size formula is prescribed.

### Prepare for the work steps

#### Setup: Define the workflow
- Examples of started and completed work

#### Policies: Agree on WIP and policies
- Defined workflow
- Current started work

#### Active work: Manage started work
- Current board
- Age and blockers of open items

#### Observe: Understand flow data
- Stable start and finish definitions
- Dated work items

#### Improve: Try one improvement
- Flow observations
- Participant feedback


## Manage the workflow together

### Ready

...

### In progress
WIP limit: ...

...

### Review
WIP limit: ...

...

### Done

...

### Working policies

**WIP boundary**
...

**Pull and completion**
...

**Blocker**
...

**SLE and learning**
...

[The Kanban Guide · original educational rendering](https://kanbanguides.org/the-kanban-guide/)


## Setup: Define the workflow
Expected artifact: Definition of Workflow

Scope · items · start · finish · states

Entry:

...

- [ ] Do columns represent actual states?
- [ ] Are start and finish unambiguous?

## Policies: Agree on WIP and policies
Expected artifact: Visible working policies

WIP · pull · blockers · SLE

Entry:

...

- [ ] Do policies prevent uncontrolled starts?
- [ ] Is the SLE labeled as a forecast?

## Active work: Manage started work
Expected artifact: Actively managed work

Next action · ownership · blocker

Entry:

...

- [ ] Does blocked work still count as WIP?
- [ ] Does aging work have a next action?

## Observe: Understand flow data
Expected artifact: Flow observations

WIP · Throughput · Work Item Age · Cycle Time

Entry:

...

- [ ] Are measurement boundaries stable?
- [ ] Are open and completed items distinguished?

## Improve: Try one improvement
Expected artifact: Improvement experiment

Change · expected effect · review

Entry:

...

- [ ] Is the expected effect observable?
- [ ] Will the change be reviewed?

## Open questions and next steps

...

Method guide: https://methodatlas.meierhoff-systems.de/en/methods/kanban/run-sheet
MarkdownShow template

Hypothesis-Driven Troubleshooting: Worksheet

Set subject, evidence, roles, flow, and review.

# Hypothesis-Driven Troubleshooting: Worksheet

## Question
To be confirmed

## Desired outcome
To be confirmed

## Scope
Complete run
Time: 30–180 minutes per failure; for production-critical events run alongside safe containment.

Complete this worksheet on paper or in your own document during the work.

## Preparation for this scope

### Method setup
- **Materials:** Precise symptom, architecture and changes, hypothesis board, predictions, safe tests, and evidence log
- **Roles:** Troubleshooting lead · hypothesis owners · operator · system experts · scribe
- **Advance information:** Set symptom, affected/unaffected cases, current signals, and safe test boundaries; separate containment from diagnosis.
- **Overall time needed:** 30–180 minutes per failure; for production-critical events run alongside safe containment.
- **Setup:** Prepare a workspace with Symptom · Hypotheses · Predictions · Test · Decide; separate example from live data.

### Prepare for the work steps

#### Symptom: Symptom view
- Set symptom, affected/unaffected cases, current signals, and safe test boundaries; separate containment from diagnosis.
- Precise symptom, architecture and changes, hypothesis board, predictions, safe tests, and evidence log

#### Hypotheses: Hypothesis space
- Symptom view

#### Predictions: Prediction matrix
- Hypothesis space

#### Test: Test result
- Prediction matrix

#### Decide: Diagnostic finding
- Test result


## Hypothesis-Driven Troubleshooting · working structure

| Hypothesis | Mechanism | Prediction | Falsifier | Test | Finding | Status |
| --- | --- | --- | --- | --- | --- | --- |
| Shared pool exhausted |   |   |   |   |   |   |
| EU network fault |   |   |   |   |   |   |
| Web release blocks |   |   |   |   |   |   |

- **Shared pool exhausted:** Add evidence and rationale.
- **EU network fault:** Add evidence and rationale.
- **Web release blocks:** Add evidence and rationale.

The blank worksheet contains prompts only.

[Reference · original teaching example](https://sre.google/sre-book/effective-troubleshooting/)


## Symptom: Symptom view
Expected artifact: Symptom view

Operationalise observed failure and collect Is/Is-Not cases.

Entry:

...

- [ ] Is success/failure measurable and free of causal claims?

## Hypotheses: Hypothesis space
Expected artifact: Hypothesis space

Derive several mechanistic explanations from signals, architecture, and changes.

Entry:

...

- [ ] Does each explain affected and unaffected cases?

## Predictions: Prediction matrix
Expected artifact: Prediction matrix

Write observable predictions and falsifying signals for every hypothesis.

Entry:

...

- [ ] Are predictions recorded before testing and discriminating?

## Test: Test result
Expected artifact: Test result

Run safest highest-information test, changing one condition at a time.

Entry:

...

- [ ] Does test vary exactly the discriminating condition and remain reversible?

## Decide: Diagnostic finding
Expected artifact: Diagnostic finding

Update hypotheses against all evidence, record mechanism, and hand over next action/residual uncertainty.

Entry:

...

- [ ] Does leading hypothesis explain all signals with a serious alternative tested?

## Open questions and next steps

...

Method guide: https://methodatlas.meierhoff-systems.de/en/methods/hypothesis-driven-troubleshooting/run-sheet
MarkdownShow template

Fault Isolation: Worksheet

Set subject, evidence, roles, flow, and review.

# Fault Isolation: Worksheet

## Question
To be confirmed

## Desired outcome
To be confirmed

## Scope
Complete run
Time: 30–180 minutes per failure; duration depends on test access and system complexity.

Complete this worksheet on paper or in your own document during the work.

## Preparation for this scope

### Method setup
- **Materials:** Symptom and reproduction, architecture/dependency map, comparison signals, safe tests, and isolation log
- **Roles:** Troubleshooting lead · system/component experts · operator · scribe
- **Advance information:** Set symptom, success/failure criterion, affected and unaffected cases, and safe test boundaries.
- **Overall time needed:** 30–180 minutes per failure; duration depends on test access and system complexity.
- **Setup:** Prepare a workspace with Boundary · Candidates · Split · Test · Confirm; separate example from live data.

### Prepare for the work steps

#### Boundary: Fault boundary
- Set symptom, success/failure criterion, affected and unaffected cases, and safe test boundaries.
- Symptom and reproduction, architecture/dependency map, comparison signals, safe tests, and isolation log

#### Candidates: Candidate map
- Fault boundary

#### Split: Isolation plan
- Candidate map

#### Test: Isolation log
- Isolation plan

#### Confirm: Confirmed domain
- Isolation log


## Fault Isolation · working structure

| Test | Candidates before | Test change | Discriminating prediction | Result | Excluded | Remaining |
| --- | --- | --- | --- | --- | --- | --- |
| Direct API call |   |   |   |   |   |   |
| Isolated pool |   |   |   |   |   |   |

- **Direct API call:** Add evidence and rationale.
- **Isolated pool:** Add evidence and rationale.

The blank worksheet contains prompts only.

[Reference · original teaching example](https://sre.google/sre-book/effective-troubleshooting/)


## Boundary: Fault boundary
Expected artifact: Fault boundary

Describe reproducible symptom and collect Is/Is-Not cases.

Entry:

...

- [ ] Are affected, unaffected, and trigger concrete?

## Candidates: Candidate map
Expected artifact: Candidate map

Map dependency path and collect plausible fault domains without fixing cause.

Entry:

...

- [ ] Does map cover signal path end to end?

## Split: Isolation plan
Expected artifact: Isolation plan

Choose the most discriminating test, preferably safe bisection or affected/unaffected comparison.

Entry:

...

- [ ] Will either result exclude several candidates?

## Test: Isolation log
Expected artifact: Isolation log

Run one test at a time, record result and side effect, and update candidate space.

Entry:

...

- [ ] Was only one distinguishing condition changed?

## Confirm: Confirmed domain
Expected artifact: Confirmed domain

Confirm isolated area by reproduction or controlled return and hand over residual uncertainty.

Entry:

...

- [ ] Does symptom occur with the domain and disappear without it?

## Open questions and next steps

...

Method guide: https://methodatlas.meierhoff-systems.de/en/methods/fault-isolation/run-sheet
MarkdownShow template

Golden Path: Worksheet

Choose a common developer use case and supported target environment; inspect access, security requirements and existing tools.

# Golden Path: Worksheet

## Question
To be confirmed

## Desired outcome
To be confirmed

## Scope
Complete cycle
Time: Multiple sessions for construction and pilot; ongoing maintenance

Complete this worksheet on paper or in your own document during the work.

## Preparation for this scope

### Method setup
- **Materials:** Reference repository, runnable development environment, template, CI/CD, checklist and feedback channel.
- **Roles:** Platform team, a pilot development team, security and operations owners.
- **Advance information:** Choose a common developer use case and supported target environment; inspect access, security requirements and existing tools.
- **Overall time needed:** Multiple sessions for construction and pilot; ongoing maintenance
- **Setup:** Prepare a fresh test environment and reference repository; follow instructions alongside the running workflow.

### Prepare for the work steps

#### Standard case: Bound the standard case
- User needs
- Platform capabilities

#### Build route: Build an executable route
- Use case
- Environment and tool access

#### Operations: Include quality and operations
- Reference path
- Quality and operational requirements

#### Pilot: Pilot with a team
- Operable route
- Pilot team

#### Maintenance: Support and exceptions
- Pilot feedback
- Maintenance owner


## Golden Path· worksheet

### Supported case
Who uses the path for which outcome and what is outside scope?

...

### Prerequisites
Which access, tools and knowledge are checked before starting?

...

### Executable route
Specify steps, linked templates and expected results.

...

### Quality and operations
Which checks, operating information and recovery steps are included?

...

### Maintenance and exceptions
Who maintains which version and how are support or deviations handled?

...

Answer the prompts with your own information. Explicitly mark unknowns and assumptions.

[Spotify Engineering: How We Use Golden Paths](https://engineering.atspotify.com/2020/8/how-we-use-golden-paths-to-solve-fragmentation-in-our-software-ecosystem)


## Standard case: Bound the standard case
Expected artifact: Use case

Are audience and supported case unambiguous?

Entry:

...

- [ ] Are audience and supported case unambiguous?

## Build route: Build an executable route
Expected artifact: Reference route

Can someone follow the route without insider knowledge?

Entry:

...

- [ ] Can someone follow the route without insider knowledge?

## Operations: Include quality and operations
Expected artifact: Operating standard

Is the generated result operable?

Entry:

...

- [ ] Is the generated result operable?

## Pilot: Pilot with a team
Expected artifact: Pilot feedback

Did the pilot team reach the endpoint independently?

Entry:

...

- [ ] Did the pilot team reach the endpoint independently?

## Maintenance: Support and exceptions
Expected artifact: Supported route

Are support and the exception route visible?

Entry:

...

- [ ] Are support and the exception route visible?

## Open questions and next steps

...

Method guide: https://methodatlas.meierhoff-systems.de/en/methods/golden-path/run-sheet