methodatlas
RunsheetUX Research

Tree Testing

ComplexityMedium
Time1-2 Tage
Participants10-30
FormatAsync
MaturityCanonical
01

Prerequisite

What needs to be finished first

Complete firstIA hypothesis

A concrete IA structure (hierarchy from card sorting or existing navigation) exists and should be tested.

Without: Without a testable structure, there is no test object and findability measurement is not reproducible.
Complete firstRealistic tasksnot in catalog

5-10 typical search tasks from real user scenarios are formulated without revealing the target category in the task wording.

Without: If the task wording names the target category, the test measures reading ability instead of findability.
02

Preparation

What needs to be ready before start

Materials

Tree-testing tool (Treejack by Optimal Workshop, UserZoom, Maze); tree structure as JSON or CSV; tasks with target path marked; survey setup; recruiting link.

People / roles

One UX researcher as owner; recruiter for 15-30 participants per variant; optional stakeholder for task validation; notes recipient in the team.

Pre-read

IA structure as tree with all nodes and leaf nodes; task list with target paths; hypotheses (which tasks likely fail); target-group demographics.

Time needed

1-2 days setup, 3-7 days data collection, 1-2 days analysis

Setup

Build tree in the tool (all nodes exactly as planned). Formulate tasks: concrete search task, do NOT name category. Pilot with 2-3 people. Main run with 15-30 participants (more for variant comparison).

03

Core question

The one question this method answers

Can participants find the correct answers in the proposed navigation structure, and which paths lead them astray?

04

Flow

Marker: Phase

StepDurationActionHint
1Phase 1: Build tree
2-4 hImport or manually create IA structure in the tool. Label every node exactly as planned. Omit visual aids, only text hierarchy.One wrong label in the tree distorts all downstream tasks. Before pilot, check tree against IA document node by node.
2Phase 2: Formulate tasks
2-3 hFormulate realistic search scenario per task. Mark target path in the tool (can have several correct paths).Avoid category words in task text. If category is "Account & Profile", do not write "in your account" in the task, otherwise the text gives away the answer.
3Phase 3: Pilot
1-2 hLet 2-3 people run through the test. Watch for confusion: unclear tasks, ambiguous labels, tree too deep. Adjust tasks or tree after pilot.Pilot reveals format errors early. Running the main test directly after a faulty pilot wastes data.
4Phase 4: Main run
3-7 days runtime15-30 participants per tree variant. For A/B comparison (old vs new), at least 30 per group. Anonymous asynchronous work, tool records paths.Below 10 participants, statistical signal is thin. Multiple variants require exponentially more participants.
5Phase 5: Analysis and iteration
1-2 daysPer task: success rate (correct leaf nodes), directness (direct path without detours), most frequent wrong paths. Tasks below 60% success rate are critical. Adjust label or hierarchy, retest if needed.Success rate alone misleads. Low directness plus high success means trial and error. Path analysis shows where users fail.
05

Artifact

What comes out at the end

Form

Tree-testing report with tree visualization, task list, success-rate table per task, directness values, top-3 wrong paths per task, sample description, and prioritized recommendations for IA adjustment.

Versioning / ownership

Own test run per tree version with date. Explicitly document comparison between versions (improvement per task in percentage points). Store tree snapshots as JSON export in repo so IA history remains traceable.

Tool alternatives
  • Treejack by Optimal Workshop
  • UserZoom with tree-test module
  • Maze for structural tests
  • Lyssna (formerly UsabilityHub)
  • Custom setup with survey tool plus manual analysis

tree-testing-working-template.md

Compact working template for Tree Testing with context, input, output artifacts, and next step.

Tree Testing Canvas

Context

What is this method used for?

Core question

Which question should be answered at the end?

Input

Which data, observations, or materials are available?

Working area

  • Area 1:
  • Area 2:
  • Area 3:
  • Relationships / patterns:

Output artifacts

  • Findability Metrics:
  • Path Analysis:
  • Revised IA:

Open questions

  • ...

Next step

Owner, date, success signal.

06

Example output

Concrete filled scenario, fictional example

tree-testing-beispiel.md

Concrete filled scenario, fictional example

Tree Testing - Help Center IA v2 (CW 20/2026, n=24)

Tree variant: New IA with 6 top-level categories (see Card Sorting CW 18) Sample: 24 existing users, asynchronous via Treejack, EUR 10 incentive

Task results

TaskSuccess RateDirectnessMost frequent wrong path
Reset password96% (23/24)87%Account & Security -> Account Settings -> correct
Download invoice79% (19/24)63%Billing -> Invoices correct (error: 5 first went to Features)
Create API token42% (10/24)28%CRITICAL: 8 went to Account, 6 to Features, only 10 to Integrations & API
Invite employee71% (17/24)54%Account -> Employees wrong (correct: Getting Started)
Cancel subscription88% (21/24)79%Billing -> Cancellation correct

Insights

  • API-token task is show-stopper: 58% fail, almost no direct path usage. Label "Integrations & API" is not recognized as API management.
  • Invite employee is searched in account area (mental model), not Getting Started. Consider moving or cross-linking.
  • Invoice download works, but 21% first go to Features. Check clarity of "Features" label.

Recommendations

  1. Rename "Integrations & API" to "Developers & API" and increase top-level visibility
  2. Cross-link "Invite employee" additionally from Account
  3. Rename or split Features area more clearly
07

Pitfalls

Recognize symptoms and steer against them

Trap

Task text gives away answer

Symptom

Task contains category keyword, testers find path too easily.

What to do

Formulate task without category language. Instead of "Where do you find billing?", use "You want to download the last invoice, where do you click?" Pilot checks wording.

Trap

Too few participants

Symptom

Test with 5-8 participants, success rates fluctuate extremely.

What to do

At least 15 for pattern insight, 30+ for statistically more reliable comparisons. A/B needs separate samples.

Trap

Tree too deep

Symptom

Hierarchy has 5+ levels, testers lose orientation and jump back.

What to do

Limit depth to max 3-4 levels. If more is needed, restructure IA instead of tree-testing.

Trap

Findability confused with visuals

Symptom

Stakeholders expect insight about design, tree test only provides text-structure insight.

What to do

Set expectations upfront: tree test measures IA and labels, not visual design. For design feedback, run usability test or first-click test.

Trap

No iteration after test

Symptom

Test shows problems, IA is rolled out unchanged.

What to do

Use test findings as gate for IA launch. Below 60% success rate, iterate and retest instead of launching.

08

Stop criteria

Done signals checkable in under a minute

IA structure is not stable or changes during the test.
No 15+ target-group participants can be recruited.
Tasks cannot be formulated without revealing the category, IA labels are too specific.
Tool setup measures only end nodes, not paths, so directness analysis is missing.
Stakeholder rejects text-based test and wants to test live prototype.
Tree has fewer than 15 nodes, manual structure review is enough.

Finished the runsheet?

Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.