methodatlas
RunsheetAgile

T-Shirt Sizing

ComplexityLow
Time15-45 min
Participants2-12
FormatWorkshop + async
MaturityEstablished
01

Prerequisite

What needs to be finished first

Complete firstProduct Vision Board

A product direction or roadmap idea with themes or epics exists, so items to size are named at all.

Without: Without topic context, team sizes randomly assembled ideas that later do not fit roadmap context.
02

Preparation

What needs to be ready before start

Materials

Whiteboard or Miro board with 5 columns (XS, S, M, L, XL); sticky notes or cards with one item each; two to three predefined reference items per size; timer; pen colors for markings.

People / roles

One facilitator who guides sorting; one Product Owner for item context; implementing team (engineering, design depending on item type); optional tech lead for architecture assessment.

Pre-read

List of epics or items to size with short description; list of reference items from past releases with their real size; known constraints (team capacity, external dependencies).

Time needed

15-45 min

Setup

Five columns on wall: XS (<1 sprint), S (1 sprint), M (2-3 sprints), L (quarter), XL (half-year+). Put reference items into every column. State rule: no hour or point discussion, only comparison with references.

03

Core question

The one question this method answers

Which size class does each item fall into relative to our references, and which items are so large that they must be cut before planning?

04

Flow

Marker: Phase

StepDurationActionHint
1Phase 1: Calibrate references
10 minReview reference items per size class together. Anyone with different understanding says so now. Re-sort references if needed until group has consensus.Without calibrated references, sorting is gut feeling. If no references from real releases exist, select some from roadmap history before workshop.
2Phase 2: Silent sorting
15 minItems on wall in random order. Participants silently sort into columns, may move items placed by others. No discussion in this phase.If item moves back and forth between two people several times, it is a discussion candidate for phase 3. Mark items silently with dot.
3Phase 3: Clarify contested items
15 minMax 2 min per contested item: highest and lowest estimate explain assumptions, then decision. Mark XL items as slicing candidates.With more than 5 contested items, group either lacks context or references do not fit. Pause workshop, add discovery.
4Phase 4: Split or park XL items
10 minFor each XL item decide: split (which smaller items it breaks into), spike (discovery first) or park (does not currently fit). Document result in roadmap backlog.XL items without split remain wishful thinking. If no split proposal possible, solution understanding missing, item belongs in discovery.
05

Artifact

What comes out at the end

Form

Roadmap table or board with columns per size, every item assigned, plus separate section for slicing candidates and spike needs. Assumptions per item as short note.

Versioning / ownership

New version per roadmap review (usually quarterly). Document size changes with date and rationale in item, do not overwrite old values. On item split: mark old ID as "resolved into X, Y, Z".

Tool alternatives
  • Miro or FigJam with T-Shirt Sizing template
  • Notion database with size field
  • Jira or Linear with custom label
  • ProductBoard with size field
  • Google Sheet with size column and sorting

t-shirt-sizing-working-template.md

Compact working template for T-Shirt Sizing with context, input, output artifacts, and next step.

T-Shirt Sizing Canvas

Context

What is this method used for?

Core question

Which question should be answered at the end?

Input

Which data, observations, or materials are available?

Working area

  • Area 1:
  • Area 2:
  • Area 3:
  • Relationships / patterns:

Output artifacts

  • Size Buckets:
  • Rough Backlog Map:
  • Split Candidates:

Open questions

  • ...

Next step

Owner, date, success signal.

06

Example output

Concrete filled scenario, fictional example

t-shirt-sizing-beispiel.md

Concrete filled scenario, fictional example

T-Shirt Sizing — Roadmap H2/2026, 2026-05-18

References

  • XS: Onboarding tooltip update (3 days)
  • S: Single Sign-On integration with Google
  • M: Client export as CSV/JSON
  • L: Multi-tenant client management (completed Q1 2026)
  • XL: Complete migration to new database (parked since 2025)

Result (12 items)

  • XS: Revise tooltips
  • S: DATEV import validation, recommendation prompt in dashboard
  • M: API webhook system, new reporting modules
  • L: Client self-service portal (with @ben), workflow builder
  • XL: AI-supported receipt recognition (slicing required), white-label variant

Slicing/Spike

  • AI receipt recognition -> spike "provider comparison" (2 weeks, @anna).
  • White label -> three substories proposed for Q4: branding settings, custom domain, client whitelabel toggle.
07

Pitfalls

Recognize symptoms and steer against them

Trap

Sizes converted to hours

Symptom

Discussion revolves around "M is about 4 weeks, right?" instead of comparison with reference item.

What to do

Facilitator interrupts: T-shirt sizes are comparison classes, not hour values. Anyone needing forecast uses velocity or Monte Carlo separately.

Trap

Missing references

Symptom

Columns are empty, group sorts by feeling without calibration anchor.

What to do

Pause workshop, find three to five reference items from real delivery history and sort them. Without references, method is gut feeling with class label.

Trap

Everything becomes L or XL

Symptom

Majority of items lands in large columns, small columns empty.

What to do

Either roadmap too ambitious or items cut too coarsely at epic level. Bring to medium granularity before sizing or sharpen comparison scale.

Trap

Item content unclear

Symptom

Participants repeatedly ask what an item means, sorting stalls.

What to do

Sizing workshop is not discovery. Sort unclear items out and send to separate refinement session. Size only understood items.

Trap

PO sizes alone

Symptom

Product Owner distributes sizes without implementation team, engineering corrects drastically later.

What to do

At least one tech lead per item area present. Sizing is team estimate. PO brings scope, team brings effort assessment.

08

Stop criteria

Done signals checkable in under a minute

No reference items from real delivery history available, scale is not calibrated.
Items are in different granularity (tasks to epics mixed), comparison not meaningful.
No implementation team present, only stakeholders or PO, sizes become wishful thinking.
More than half of items need discovery first, sizing would be speculation.
Roadmap context entirely missing, items have no relationship to strategic themes.
Workshop cut under 15 min, with more than 8 items not enough time for contested cases.

Finished the runsheet?

Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.