methodatlas
RunsheetProduct Discovery

Hypothesis Prioritization Canvas

ComplexityMedium
Time60-90 min
Participants3-8
FormatWorkshop
MaturityEstablished
01

Prerequisite

What needs to be finished first

Complete firstAssumption Mapping

A list of hypotheses or assumptions with initiative context exists (typically 5-20 entries).

Without: Without a hypothesis list, the canvas becomes a brainstorming session and prioritization value is lost.
02

Preparation

What needs to be ready before start

Materials

Canvas with three axes (risk, evidence, effort) or a quadrant; hypothesis list as cards; visible scale definitions; pens; timer; voting tool (Miro, FigJam).

People / roles

A facilitator who holds scales and moderates consensus; one Product Lead as hypothesis owner; 3-6 participants from Discovery, Engineering, Design, Data; a scribe for ratings.

Pre-read

Hypothesis list distributed 1-2 days before; scale definitions (risk 1-5, evidence 1-5, effort 1-5); existing evidence from analytics, interviews, or prior tests; test tool options with rough effort estimates.

Time needed

60-90 min

Setup

Canvas with three axes or quadrants. Scale banners on the wall: risk (consequence if hypothesis is wrong), evidence (degree of support), effort (test cost). Keep hypothesis cards ready.

03

Core question

The one question this method answers

Which hypotheses combine high risk, low evidence, and low testing effort, and which of them should be tested in the next iterations?

04

Flow

Marker: Phase

StepDurationActionHint
1Phase 1: Calibrate scales
10 minReview scale definitions with examples. At least one example per level (for example, "Risk 5: business model at stake without confirmation"). Record consensus.Without scale calibration, rating falls back to gut feel. If no shared scale definition exists, the team has no baseline for prioritization.
2Phase 2: Rate hypotheses individually
30-40 minAssign three values per hypothesis: risk, evidence, effort. Vote anonymously first, then reach consensus. If disagreement occurs, ask for underlying assumptions, do not enforce consensus.Max 3 min per hypothesis. If discussion is long, document the assumption and park it for follow-up. Hypothesis evaluation is triage, not deep analysis.
3Phase 3: Position on canvas
10 minPlace hypotheses on the canvas. Top-right quadrant: high risk, low evidence, low effort. These are top test candidates.If top-right is empty, hypotheses are either too generic or tests are too large. Re-scope hypotheses or split tests.
4Phase 4: Test backlog and owners
10-20 minSelect the top 3 hypotheses for next tests. Set owner and start date per hypothesis. Document follow-up hypotheses as backlog. Fix pipeline order.Maximum three parallel tests for small teams. For discovery teams with more than five people, more tests may be possible. Pipeline discipline: do not start everything at once.
05

Artifact

What comes out at the end

Form

Prioritization canvas as board export, plus a hypothesis backlog in Markdown or table with ratings (R/E/A), canvas position, owner, test status, planned start date.

Versioning / ownership

Snapshot per discovery sprint with date. Track hypotheses with status (Planned, Tested, Validated, Rejected). On status change, add date and test reference. Keep backlog pipeline updated and archive prior versions.

Tool alternatives
  • Strategyzer Hypothesis Prioritization template
  • Miro or FigJam with canvas template
  • Notion database with risk/evidence/effort properties
  • Productboard with custom fields
  • Confluence page with table and embedded canvas

hypothesis-prioritization-canvas-working-template.md

Compact working template for Hypothesis Prioritization Canvas with context, input, output artifacts, and next step.

Hypothesis Prioritization Canvas Canvas

Context

What is this method used for?

Core question

Which question should be answered at the end?

Input

Which data, observations, or materials are available?

Working area

  • Area 1:
  • Area 2:
  • Area 3:
  • Relationships / patterns:

Output artifacts

  • Prioritization Canvas:
  • Hypothesis backlog:

Open questions

  • ...

Next step

Owner, date, success signal.

06

Example output

Concrete filled scenario, fictional example

hypothesis-prioritization-canvas-beispiel.md

Concrete filled scenario, fictional example

Hypothesis Prioritization - Invoice pre-classification, 2026-05-18

Scales

  • Risk: 1=cosmetic, 3=iteration delay, 5=business-model critical.
  • Evidence: 1=no data, 3=anecdotal, 5=quantitative validation.
  • Effort: 1=Survey 1 day, 3=landing page 1 week, 5=prototype > 2 weeks.

Hypothesis pipeline (8 rated)

IDHypothesisREAPosition
H1Solo tax advisors pay >EUR 29 per month512Top-right (Test 1)
H2AI hit rate of >80% is acceptable413Top-right (Test 2)
H3LinkedIn ads reach target group322Right (Test 3)
H4DATEV interface without custom adapter432Mid
H5Most receipts are PDFs241Bottom-left (rejected, evidence sufficient)
H6Advisors use mobile-first212Mid-left (backlog)
H7Recommendation mode drives growth314Mid (after H1-H3)
H8Compliance requirements for AI processing are solvable524Mid (parallel clarification)

Top-3 for test pipeline

  1. H1 pricing willingness - landing page test, Owner: @lisa, start 20.05.
  2. H2 AI hit rate - concierge test with 10 users, Owner: @anna, start 03.06.
  3. H3 distribution - ad campaign, Owner: @marcus, start 27.05.
07

Pitfalls

Recognize symptoms and steer against them

Trap

Scales without examples

Symptom

Ratings are inconsistent, and risk 4 for one hypothesis equals risk 2 for another.

What to do

Before rating, define examples per scale level in initiative context. Keep scale definitions written at the wall. If unclear, pause and recalibrate the scale.

Trap

Evidence becomes gut feel

Symptom

High evidence is assigned without sources; hypotheses are prematurely marked as "validated."

What to do

For evidence >3, require a source (study, interview quote, analytics extract). Without source, evidence is 1 or 2. Do not reject hypotheses, only validate via evidence.

Trap

Risk and effort mixed

Symptom

Rating "high risk, high effort" is handled as "do not test," and high-risk hypotheses are ignored.

What to do

Keep three axes separate. Split high-risk and high-effort hypotheses into sub-hypotheses or pilot with a smaller test design.

Trap

Too many parallel tests

Symptom

All top 5 hypotheses are started at once and capacity is insufficient, so tests are executed poorly.

What to do

Run a maximum of 3 parallel tests in a standard discovery team. Keep pipeline sequential or staged. Better few high-quality tests than many half-baked ones.

Trap

Hypothesis drop-off

Symptom

After the workshop the canvas is forgotten, hypotheses disappear into a Confluence page without follow-up.

What to do

Maintain pipeline backlog in a central tool (Productboard, Notion). Keep test status current. Review weekly with canvas updates.

08

Stop criteria

Done signals checkable in under a minute

Fewer than three hypotheses and prioritization becomes overhead.
Scales cannot be defined clearly, so ratings are purely gut feel.
No discovery capacity is available for testing, so prioritization leads nowhere.
The initiative is already in build phase and test prioritization is not the next step.
Stakeholder wants only one predetermined hypothesis tested; prioritization would become window dressing.
Hypotheses are not falsifiable (too vague), so rating is meaningless.

Finished the runsheet?

Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.