At least 10 comparable ideas or experiments are available in a list with short hypothesis statements.
ICE Scoring
Prerequisite
What needs to be finished first
Preparation
What needs to be ready before start
Table (idea, hypothesis, I, C, E, score, owner, status); visible scale definitions; example values for 1, 5, and 10 per dimension.
A facilitator (Growth Lead or PM); 3-6 raters from Product, Engineering, Marketing, or Data; a judge when variance is high.
Current growth goals or outcome context; previous experiments with results for confidence calibration; known capacity per sprint.
30-60 min
Share table, prepare columns I, C, E with 1-10 scale. Set score anchors before start: what is Impact 10 in your context (for example move North Star by 5%), what is Ease 10 (for example <2 days of development).
Core question
The one question this method answers
Which ideas give us the highest learning rate per effort and which should we intentionally exclude?
Flow
Marker: Phase
| Step | Duration | Action | Hint |
|---|---|---|---|
1Phase 1: Calibrate scales | 10 min | Use three recent experiments as references: one high impact, one low impact, one medium. Make values explicit so each participant applies the same scale. | If no reference experiments exist, anchor to concrete values from the backlog. Otherwise everyone scales against personal gut feel. |
2Phase 2: Solo rating | 15-20 min | Each rater assigns I, C, and E per idea alone, without discussion. Do not compute averages; keep individual points visible. | Discussion in this phase distorts distribution. Reveal ratings only after all solo inputs are in. |
3Phase 3: Check spread | 15 min | Inspect spread per idea. If a dimension differs by more than 3 points, briefly discuss what each rater sees differently. Keep assumptions explicit; do not force averages. | High spread is information, not a defect. It can indicate different assumptions or knowledge levels. |
4Phase 4: Score and cut-off | 10-15 min | Calculate score per idea (I * C * E or average of raters). Set a cut-off, for example top N by score or above a threshold. Allow justified exceptions. | Multiplication amplifies extreme values. If an idea scores below 3 in any dimension, it should usually be excluded rather than rescued by other high dimensions. |
Artifact
What comes out at the end
Table with idea, hypothesis, individual ratings, aggregate score, status (top, backlog, rejected), owner, and date. Plus short notes on scale anchors and spread discussion.
Include date and raters in the header each scoring round. Do not overwrite old scores; add a new column or snapshot so confidence calibration can be learned over time.
- Google Sheet with formulas and sort function
- Notion or Coda database with filters
- Productboard or Reveall for idea management
- Linear issue list with ICE properties
ice-scoring-working-template.md
Compact working template for ICE Scoring with context, input, output artifacts, and next step.
ICE Scoring Working Template
Goal
Prioritizes ideas by Impact, Confidence, and Ease.
Context
When and for what do we use this method?
Input
Which data, observations, decisions, or materials are available?
Execution
Short notes along the runsheet.
Output artifacts
- ICE table:
- Top ideas list:
Assumptions and open questions
- ...
Decision / next step
Owner, date, and success signal.
Example output
Concrete filled scenario, fictional example
ice-scoring-beispiel.md
Concrete filled scenario, fictional example
ICE Scoring — Activation Squad, KW 21/2026
Scale anchors: Impact 10 = +5% D7 retention, Ease 10 = implementable in under 2 days.
| Idea | I | C | E | Score | Status |
|---|---|---|---|---|---|
| Tooltip in empty workspace | 7 | 8 | 9 | 504 | Top |
| Recommendation email at day 3 | 8 | 5 | 6 | 240 | Top |
| Re-engagement push 30 days | 6 | 4 | 7 | 168 | Backlog |
| Verification flow rewrite | 9 | 6 | 2 | 108 | Backlog |
| Personal quiz onboarding | 4 | 3 | 5 | 60 | Rejected |
Spread note: Idea "verification flow" scored Ease 1 by @ben, Ease 4 by @anna (different assumptions about auth refactor implications).
Pitfalls
Recognize symptoms and steer against them
Scales without anchors
Raters submit values between 6 and 9 for all ideas, and spread is lacking.
Before the next round, set three real anchors across the full scale: one at low, one at medium, one at high. Ask raters to evaluate against these anchors.
Confidence as optimism
Confidence values are mostly 8-10 because no one wants to appear uncertain.
Ground confidence in prior success or data. Confidence 9 requires at least one comparable positive experiment in existing portfolio.
Averages without discussion
Raters average values immediately, and divergent assumptions disappear.
Expose spread first, then discuss. A difference of more than 3 points is the trigger for assumption clarification.
Score as truth
Team treats top-N as an order, without checking strategic context.
Cross-check the top list against current outcomes or OKRs. Deviations are allowed, but document them in the log.
Scale drift over time
After three rounds, Impact 8 and Ease 7 mean different things than at the first round.
Reconfirm or adapt anchor examples every 4-6 weeks with the team. Document drift, do not let it accumulate silently.
Misusing ICE for strategy
Quarterly strategy or large platform investments are prioritized with ICE.
Switch to the right tool: WSJF, CoD, or qualitative strategy methods. ICE should stay in the experiment backlog.
Stop criteria
Done signals checkable in under a minute
Finished the runsheet?
Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.