A list of hypotheses or assumptions with initiative context exists (typically 5-20 entries).
Hypothesis Prioritization Canvas
Prerequisite
What needs to be finished first
Preparation
What needs to be ready before start
Canvas with three axes (risk, evidence, effort) or a quadrant; hypothesis list as cards; visible scale definitions; pens; timer; voting tool (Miro, FigJam).
A facilitator who holds scales and moderates consensus; one Product Lead as hypothesis owner; 3-6 participants from Discovery, Engineering, Design, Data; a scribe for ratings.
Hypothesis list distributed 1-2 days before; scale definitions (risk 1-5, evidence 1-5, effort 1-5); existing evidence from analytics, interviews, or prior tests; test tool options with rough effort estimates.
60-90 min
Canvas with three axes or quadrants. Scale banners on the wall: risk (consequence if hypothesis is wrong), evidence (degree of support), effort (test cost). Keep hypothesis cards ready.
Core question
The one question this method answers
Which hypotheses combine high risk, low evidence, and low testing effort, and which of them should be tested in the next iterations?
Flow
Marker: Phase
| Step | Duration | Action | Hint |
|---|---|---|---|
1Phase 1: Calibrate scales | 10 min | Review scale definitions with examples. At least one example per level (for example, "Risk 5: business model at stake without confirmation"). Record consensus. | Without scale calibration, rating falls back to gut feel. If no shared scale definition exists, the team has no baseline for prioritization. |
2Phase 2: Rate hypotheses individually | 30-40 min | Assign three values per hypothesis: risk, evidence, effort. Vote anonymously first, then reach consensus. If disagreement occurs, ask for underlying assumptions, do not enforce consensus. | Max 3 min per hypothesis. If discussion is long, document the assumption and park it for follow-up. Hypothesis evaluation is triage, not deep analysis. |
3Phase 3: Position on canvas | 10 min | Place hypotheses on the canvas. Top-right quadrant: high risk, low evidence, low effort. These are top test candidates. | If top-right is empty, hypotheses are either too generic or tests are too large. Re-scope hypotheses or split tests. |
4Phase 4: Test backlog and owners | 10-20 min | Select the top 3 hypotheses for next tests. Set owner and start date per hypothesis. Document follow-up hypotheses as backlog. Fix pipeline order. | Maximum three parallel tests for small teams. For discovery teams with more than five people, more tests may be possible. Pipeline discipline: do not start everything at once. |
Artifact
What comes out at the end
Prioritization canvas as board export, plus a hypothesis backlog in Markdown or table with ratings (R/E/A), canvas position, owner, test status, planned start date.
Snapshot per discovery sprint with date. Track hypotheses with status (Planned, Tested, Validated, Rejected). On status change, add date and test reference. Keep backlog pipeline updated and archive prior versions.
- Strategyzer Hypothesis Prioritization template
- Miro or FigJam with canvas template
- Notion database with risk/evidence/effort properties
- Productboard with custom fields
- Confluence page with table and embedded canvas
hypothesis-prioritization-canvas-working-template.md
Compact working template for Hypothesis Prioritization Canvas with context, input, output artifacts, and next step.
Hypothesis Prioritization Canvas Canvas
Context
What is this method used for?
Core question
Which question should be answered at the end?
Input
Which data, observations, or materials are available?
Working area
- Area 1:
- Area 2:
- Area 3:
- Relationships / patterns:
Output artifacts
- Prioritization Canvas:
- Hypothesis backlog:
Open questions
- ...
Next step
Owner, date, success signal.
Example output
Concrete filled scenario, fictional example
hypothesis-prioritization-canvas-beispiel.md
Concrete filled scenario, fictional example
Hypothesis Prioritization - Invoice pre-classification, 2026-05-18
Scales
- Risk: 1=cosmetic, 3=iteration delay, 5=business-model critical.
- Evidence: 1=no data, 3=anecdotal, 5=quantitative validation.
- Effort: 1=Survey 1 day, 3=landing page 1 week, 5=prototype > 2 weeks.
Hypothesis pipeline (8 rated)
| ID | Hypothesis | R | E | A | Position |
|---|---|---|---|---|---|
| H1 | Solo tax advisors pay >EUR 29 per month | 5 | 1 | 2 | Top-right (Test 1) |
| H2 | AI hit rate of >80% is acceptable | 4 | 1 | 3 | Top-right (Test 2) |
| H3 | LinkedIn ads reach target group | 3 | 2 | 2 | Right (Test 3) |
| H4 | DATEV interface without custom adapter | 4 | 3 | 2 | Mid |
| H5 | Most receipts are PDFs | 2 | 4 | 1 | Bottom-left (rejected, evidence sufficient) |
| H6 | Advisors use mobile-first | 2 | 1 | 2 | Mid-left (backlog) |
| H7 | Recommendation mode drives growth | 3 | 1 | 4 | Mid (after H1-H3) |
| H8 | Compliance requirements for AI processing are solvable | 5 | 2 | 4 | Mid (parallel clarification) |
Top-3 for test pipeline
- H1 pricing willingness - landing page test, Owner: @lisa, start 20.05.
- H2 AI hit rate - concierge test with 10 users, Owner: @anna, start 03.06.
- H3 distribution - ad campaign, Owner: @marcus, start 27.05.
Pitfalls
Recognize symptoms and steer against them
Scales without examples
Ratings are inconsistent, and risk 4 for one hypothesis equals risk 2 for another.
Before rating, define examples per scale level in initiative context. Keep scale definitions written at the wall. If unclear, pause and recalibrate the scale.
Evidence becomes gut feel
High evidence is assigned without sources; hypotheses are prematurely marked as "validated."
For evidence >3, require a source (study, interview quote, analytics extract). Without source, evidence is 1 or 2. Do not reject hypotheses, only validate via evidence.
Risk and effort mixed
Rating "high risk, high effort" is handled as "do not test," and high-risk hypotheses are ignored.
Keep three axes separate. Split high-risk and high-effort hypotheses into sub-hypotheses or pilot with a smaller test design.
Too many parallel tests
All top 5 hypotheses are started at once and capacity is insufficient, so tests are executed poorly.
Run a maximum of 3 parallel tests in a standard discovery team. Keep pipeline sequential or staged. Better few high-quality tests than many half-baked ones.
Hypothesis drop-off
After the workshop the canvas is forgotten, hypotheses disappear into a Confluence page without follow-up.
Maintain pipeline backlog in a central tool (Productboard, Notion). Keep test status current. Review weekly with canvas updates.
Stop criteria
Done signals checkable in under a minute
Finished the runsheet?
Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.