A product direction or roadmap idea with themes or epics exists, so items to size are named at all.
T-Shirt Sizing
Prerequisite
What needs to be finished first
Preparation
What needs to be ready before start
Whiteboard or Miro board with 5 columns (XS, S, M, L, XL); sticky notes or cards with one item each; two to three predefined reference items per size; timer; pen colors for markings.
One facilitator who guides sorting; one Product Owner for item context; implementing team (engineering, design depending on item type); optional tech lead for architecture assessment.
List of epics or items to size with short description; list of reference items from past releases with their real size; known constraints (team capacity, external dependencies).
15-45 min
Five columns on wall: XS (<1 sprint), S (1 sprint), M (2-3 sprints), L (quarter), XL (half-year+). Put reference items into every column. State rule: no hour or point discussion, only comparison with references.
Core question
The one question this method answers
Which size class does each item fall into relative to our references, and which items are so large that they must be cut before planning?
Flow
Marker: Phase
| Step | Duration | Action | Hint |
|---|---|---|---|
1Phase 1: Calibrate references | 10 min | Review reference items per size class together. Anyone with different understanding says so now. Re-sort references if needed until group has consensus. | Without calibrated references, sorting is gut feeling. If no references from real releases exist, select some from roadmap history before workshop. |
2Phase 2: Silent sorting | 15 min | Items on wall in random order. Participants silently sort into columns, may move items placed by others. No discussion in this phase. | If item moves back and forth between two people several times, it is a discussion candidate for phase 3. Mark items silently with dot. |
3Phase 3: Clarify contested items | 15 min | Max 2 min per contested item: highest and lowest estimate explain assumptions, then decision. Mark XL items as slicing candidates. | With more than 5 contested items, group either lacks context or references do not fit. Pause workshop, add discovery. |
4Phase 4: Split or park XL items | 10 min | For each XL item decide: split (which smaller items it breaks into), spike (discovery first) or park (does not currently fit). Document result in roadmap backlog. | XL items without split remain wishful thinking. If no split proposal possible, solution understanding missing, item belongs in discovery. |
Artifact
What comes out at the end
Roadmap table or board with columns per size, every item assigned, plus separate section for slicing candidates and spike needs. Assumptions per item as short note.
New version per roadmap review (usually quarterly). Document size changes with date and rationale in item, do not overwrite old values. On item split: mark old ID as "resolved into X, Y, Z".
- Miro or FigJam with T-Shirt Sizing template
- Notion database with size field
- Jira or Linear with custom label
- ProductBoard with size field
- Google Sheet with size column and sorting
t-shirt-sizing-working-template.md
Compact working template for T-Shirt Sizing with context, input, output artifacts, and next step.
T-Shirt Sizing Canvas
Context
What is this method used for?
Core question
Which question should be answered at the end?
Input
Which data, observations, or materials are available?
Working area
- Area 1:
- Area 2:
- Area 3:
- Relationships / patterns:
Output artifacts
- Size Buckets:
- Rough Backlog Map:
- Split Candidates:
Open questions
- ...
Next step
Owner, date, success signal.
Example output
Concrete filled scenario, fictional example
t-shirt-sizing-beispiel.md
Concrete filled scenario, fictional example
T-Shirt Sizing — Roadmap H2/2026, 2026-05-18
References
- XS: Onboarding tooltip update (3 days)
- S: Single Sign-On integration with Google
- M: Client export as CSV/JSON
- L: Multi-tenant client management (completed Q1 2026)
- XL: Complete migration to new database (parked since 2025)
Result (12 items)
- XS: Revise tooltips
- S: DATEV import validation, recommendation prompt in dashboard
- M: API webhook system, new reporting modules
- L: Client self-service portal (with @ben), workflow builder
- XL: AI-supported receipt recognition (slicing required), white-label variant
Slicing/Spike
- AI receipt recognition -> spike "provider comparison" (2 weeks, @anna).
- White label -> three substories proposed for Q4: branding settings, custom domain, client whitelabel toggle.
Pitfalls
Recognize symptoms and steer against them
Sizes converted to hours
Discussion revolves around "M is about 4 weeks, right?" instead of comparison with reference item.
Facilitator interrupts: T-shirt sizes are comparison classes, not hour values. Anyone needing forecast uses velocity or Monte Carlo separately.
Missing references
Columns are empty, group sorts by feeling without calibration anchor.
Pause workshop, find three to five reference items from real delivery history and sort them. Without references, method is gut feeling with class label.
Everything becomes L or XL
Majority of items lands in large columns, small columns empty.
Either roadmap too ambitious or items cut too coarsely at epic level. Bring to medium granularity before sizing or sharpen comparison scale.
Item content unclear
Participants repeatedly ask what an item means, sorting stalls.
Sizing workshop is not discovery. Sort unclear items out and send to separate refinement session. Size only understood items.
PO sizes alone
Product Owner distributes sizes without implementation team, engineering corrects drastically later.
At least one tech lead per item area present. Sizing is team estimate. PO brings scope, team brings effort assessment.
Stop criteria
Done signals checkable in under a minute
Finished the runsheet?
Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.