methodatlas
RunsheetProduct Strategy

ICE Scoring

ComplexityLow
Time30-60 min
Participants2-8
FormatWorkshop + async
MaturityEstablished
01

Prerequisite

What needs to be finished first

Complete firstIdea backlognot in catalog

At least 10 comparable ideas or experiments are available in a list with short hypothesis statements.

Without: With fewer than 10 ideas, scale overhead is not worthwhile and a direct pro/con discussion is faster.
02

Preparation

What needs to be ready before start

Materials

Table (idea, hypothesis, I, C, E, score, owner, status); visible scale definitions; example values for 1, 5, and 10 per dimension.

People / roles

A facilitator (Growth Lead or PM); 3-6 raters from Product, Engineering, Marketing, or Data; a judge when variance is high.

Pre-read

Current growth goals or outcome context; previous experiments with results for confidence calibration; known capacity per sprint.

Time needed

30-60 min

Setup

Share table, prepare columns I, C, E with 1-10 scale. Set score anchors before start: what is Impact 10 in your context (for example move North Star by 5%), what is Ease 10 (for example <2 days of development).

03

Core question

The one question this method answers

Which ideas give us the highest learning rate per effort and which should we intentionally exclude?

04

Flow

Marker: Phase

StepDurationActionHint
1Phase 1: Calibrate scales
10 minUse three recent experiments as references: one high impact, one low impact, one medium. Make values explicit so each participant applies the same scale.If no reference experiments exist, anchor to concrete values from the backlog. Otherwise everyone scales against personal gut feel.
2Phase 2: Solo rating
15-20 minEach rater assigns I, C, and E per idea alone, without discussion. Do not compute averages; keep individual points visible.Discussion in this phase distorts distribution. Reveal ratings only after all solo inputs are in.
3Phase 3: Check spread
15 minInspect spread per idea. If a dimension differs by more than 3 points, briefly discuss what each rater sees differently. Keep assumptions explicit; do not force averages.High spread is information, not a defect. It can indicate different assumptions or knowledge levels.
4Phase 4: Score and cut-off
10-15 minCalculate score per idea (I * C * E or average of raters). Set a cut-off, for example top N by score or above a threshold. Allow justified exceptions.Multiplication amplifies extreme values. If an idea scores below 3 in any dimension, it should usually be excluded rather than rescued by other high dimensions.
05

Artifact

What comes out at the end

Form

Table with idea, hypothesis, individual ratings, aggregate score, status (top, backlog, rejected), owner, and date. Plus short notes on scale anchors and spread discussion.

Versioning / ownership

Include date and raters in the header each scoring round. Do not overwrite old scores; add a new column or snapshot so confidence calibration can be learned over time.

Tool alternatives
  • Google Sheet with formulas and sort function
  • Notion or Coda database with filters
  • Productboard or Reveall for idea management
  • Linear issue list with ICE properties

ice-scoring-working-template.md

Compact working template for ICE Scoring with context, input, output artifacts, and next step.

ICE Scoring Working Template

Goal

Prioritizes ideas by Impact, Confidence, and Ease.

Context

When and for what do we use this method?

Input

Which data, observations, decisions, or materials are available?

Execution

Short notes along the runsheet.

Output artifacts

  • ICE table:
  • Top ideas list:

Assumptions and open questions

  • ...

Decision / next step

Owner, date, and success signal.

06

Example output

Concrete filled scenario, fictional example

ice-scoring-beispiel.md

Concrete filled scenario, fictional example

ICE Scoring — Activation Squad, KW 21/2026

Scale anchors: Impact 10 = +5% D7 retention, Ease 10 = implementable in under 2 days.

IdeaICEScoreStatus
Tooltip in empty workspace789504Top
Recommendation email at day 3856240Top
Re-engagement push 30 days647168Backlog
Verification flow rewrite962108Backlog
Personal quiz onboarding43560Rejected

Spread note: Idea "verification flow" scored Ease 1 by @ben, Ease 4 by @anna (different assumptions about auth refactor implications).

07

Pitfalls

Recognize symptoms and steer against them

Trap

Scales without anchors

Symptom

Raters submit values between 6 and 9 for all ideas, and spread is lacking.

What to do

Before the next round, set three real anchors across the full scale: one at low, one at medium, one at high. Ask raters to evaluate against these anchors.

Trap

Confidence as optimism

Symptom

Confidence values are mostly 8-10 because no one wants to appear uncertain.

What to do

Ground confidence in prior success or data. Confidence 9 requires at least one comparable positive experiment in existing portfolio.

Trap

Averages without discussion

Symptom

Raters average values immediately, and divergent assumptions disappear.

What to do

Expose spread first, then discuss. A difference of more than 3 points is the trigger for assumption clarification.

Trap

Score as truth

Symptom

Team treats top-N as an order, without checking strategic context.

What to do

Cross-check the top list against current outcomes or OKRs. Deviations are allowed, but document them in the log.

Trap

Scale drift over time

Symptom

After three rounds, Impact 8 and Ease 7 mean different things than at the first round.

What to do

Reconfirm or adapt anchor examples every 4-6 weeks with the team. Document drift, do not let it accumulate silently.

Trap

Misusing ICE for strategy

Symptom

Quarterly strategy or large platform investments are prioritized with ICE.

What to do

Switch to the right tool: WSJF, CoD, or qualitative strategy methods. ICE should stay in the experiment backlog.

08

Stop criteria

Done signals checkable in under a minute

Fewer than 10 ideas in the backlog, direct discussion is faster.
Scale anchors cannot be defined, so all ratings remain gut feel.
Raters lack both data and topic experience, and confidence is uniformly high.
A high-investment strategic decision is pending, so a lightweight heuristic is inappropriate.
Backlog contains mostly compliance or maintenance topics and ICE logic does not fit.
Scale drift since the last round is unresolved, so new scores are not comparable.

Finished the runsheet?

Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.