User Research

How to run a tree test for information architecture

A tree test checks whether people can find things in your navigation before you spend a sprint building it. Here is how to run one and read the numbers that matter.

CleverX Team ·
How to run a tree test for information architecture

How to run a tree test for information architecture

A tree test measures whether people can find things in your navigation by asking them to complete realistic find-it tasks using only your site’s text-based hierarchy, with no visual design, no search bar, and no page content to lean on. If most participants reach the right place quickly and without backtracking, your information architecture works. If they wander, you have found a labeling or grouping problem while it is still cheap to fix.

Information architecture is the invisible skeleton of a product. When it is right, nobody notices. When it is wrong, users cannot find the report, the setting, or the feature, and they blame themselves or churn. A tree test isolates that skeleton from everything else so you can evaluate the structure on its own, before a single screen is designed or a line of code is written. This guide covers what tree testing measures, when to reach for it, how it differs from card sorting, how to write good tasks, how many participants you need, and how to read the results.

What a tree test actually measures

A tree test presents your navigation as a simple, expandable text outline. Participants read a task, then click through the tree until they believe they have found the answer, or give up. Because you strip away color, layout, imagery, and search, the only variables in play are your labels and how you have nested them. That isolation is the point. You are not testing whether the button is the right shade of blue. You are testing whether “Account settings” is where people expect to find their billing details, and whether “Reports” or “Analytics” is the word they scan for.

The method answers three practical questions:

  1. Findability. Can people locate a given item at all?
  2. Efficiency. Do they get there directly, or do they backtrack and explore dead ends first?
  3. Label clarity. Which specific words and groupings cause people to go the wrong way?

A structure can look tidy in a spreadsheet and still fail every one of these tests. Tree testing turns the debate about “what should we call this section” from an opinion contest into a measurement. For a broader view of where this method sits alongside interviews, surveys, and usability testing, see our complete walkthrough to product research methods.

Tree testing versus card sorting

These two methods are often confused because both deal with structure and grouping, but they do opposite jobs.

Card sorting is generative. You hand participants a set of content items and ask them to group and name them the way they would expect to find them. It helps you discover the categories and language your users already carry in their heads, which is useful when you are building an architecture from scratch or reorganizing a messy one.

Tree testing is evaluative. You take a structure you have already drafted and check whether people can navigate it. It does not ask how they would organize things; it asks whether your organization works.

The two fit together in a sequence:

StageMethodQuestion it answers
Shape the structureCard sortHow do users expect this content grouped and labeled?
Draft the treeInternal designGiven that input, what is our proposed hierarchy?
Validate the treeTree testCan real users find things in the hierarchy we drafted?
Confirm in contextUsability testDoes it still work once design and content are added?

You do not always need every step, but running a card sort to shape the tree and a tree test to validate it is a reliable pairing. Tree testing is the checkpoint that catches problems before they reach a designer’s screen.

When to run a tree test

Reach for a tree test whenever the navigation itself is the thing in question. Common triggers:

  • A redesign or replatform where you are rebuilding menus and want to confirm the new structure before development.
  • A merger of two products or sites, where you have to fold two content sets into one hierarchy without losing either audience.
  • A major new content area, such as adding a resources hub, a settings overhaul, or a new product line to an existing menu.
  • Analytics that show a findability problem, like a high-value feature with low discovery, heavy reliance on search, or support tickets asking where something lives.

Because a tree test needs only a text outline, you can run it at the earliest possible moment, before wireframes exist. That timing is its biggest advantage. Fixing a label in a spreadsheet costs minutes. Fixing it after launch costs a sprint and a migration. If you are weighing which research to run at which stage, our guide to turning product research into better product decisions maps methods to the decisions they inform.

Step 1: build the tree

Export or reconstruct your navigation as a hierarchy of text labels. Include the real menu labels, nested to the depth users would actually click through, typically three to four levels. Do not clean up or clarify the labels for the test; use exactly what will ship, because ambiguous wording is one of the things you are trying to catch.

Keep the tree focused. If your product is enormous, you can test a representative branch rather than the entire structure, as long as the tasks you write live inside the part you include. A tree of 50 to 200 nodes is common. Much larger than that and you should split it into focused studies.

Step 2: write the tasks

Tasks are where tree tests succeed or fail. A good task describes a realistic goal in the user’s language and never uses the words that appear in the tree, because if the task says “find your billing settings” and the menu says “Billing,” you are testing reading, not findability.

Guidelines for writing tasks:

  • Frame a goal, not a location. Write “You want to stop receiving weekly summary emails. Where would you go?” rather than “Find the notification settings.”
  • Avoid matching vocabulary. Never echo the label. Describe the situation and let the participant map it to your words.
  • Keep each task to a single answer. Ambiguous tasks with several defensible destinations produce uninterpretable data.
  • Set the correct answer or answers. Before launch, mark which node or nodes count as success. Some tasks legitimately have two valid destinations; decide in advance.
  • Test 8 to 12 tasks per session. Fewer and you miss coverage; more and fatigue sets in and quality drops.

Prioritize tasks around the journeys that matter most: the features tied to activation, revenue, or support load. A tree test with well-chosen tasks tells you far more than one that exhaustively covers low-traffic corners.

Step 3: recruit the right participants

Tree testing is quantitative, so the numbers only mean something if the people generating them resemble your real users. Someone unfamiliar with your domain will guess, and their guesses inflate the noise in your success rates.

For sample size, plan on 30 to 50 participants per audience segment. That range gives success and directness percentages stable enough to trust and to compare across tasks. If you are A/B testing two competing trees, size each arm to at least 30. For narrow B2B roles where the addressable pool is small, 20 to 30 verified participants still delivers a usable directional read, though you should treat single-digit differences between tasks as noise.

Audience match matters more than raw volume. If your product serves security engineers, procurement leads, or clinicians, a panel of general consumers will produce success rates that look fine and mean nothing, because those participants do not carry the mental models your real users do. This is exactly where recruiting quality decides whether the study is worth running.

CleverX is a B2B research platform with more than 8 million verified professionals, alongside B2C reach, across 150 or more countries. Participants are identity and employment verified, so when you tree test an enterprise navigation with, say, IT decision makers or finance managers, the people clicking through the tree actually hold those roles. Recruitment typically delivers in about two to five days on a pay-as-you-go basis, which fits the pre-build timing where tree tests are most useful. Tree tests only earn trust when you have enough of the right users behind the success rates. For the mechanics of sourcing participants, our guide to recruiting participants for user research covers eight methods, and for specialized roles see how to recruit B2B research participants.

Step 4: run the study and read the metrics

Tree tests run unmoderated, which is part of why they scale well. Participants complete the tasks on their own through a testing tool, and you review aggregate results. Here is what the core metrics tell you.

MetricWhat it tells you
Success rateShare of participants who reached a correct answer. Your headline findability number for each task.
DirectnessShare who reached the answer without backtracking up the tree. Measures how obvious the path was.
Time to completeHow long the task took. Long times often flag hesitation even when success is high.
First clickWhere participants went first. Reveals which top-level label they associate with the task.
Path analysisThe routes people actually took, including dead ends. Pinpoints exactly where the structure misleads.

The two numbers to read together are success rate and directness. A task with 90 percent success but 40 percent directness is a warning, not a win: most people eventually found it, but only after wandering, which means the structure fought them and they brute-forced their way through. In a real interface with more distractions, many of those people would have given up. Aim for high success and high directness on your priority tasks.

First-click data is often the richest diagnostic. If a task has low success and most participants’ first click went to the wrong top-level section, you have located the exact label that misleads them, and you know which two sections are competing in users’ minds. That is a precise, fixable finding.

Step 5: act on the results

Findings only matter if they change the tree. Work task by task, from the lowest success rates up.

  • Low success, wrong first click clustered on one label. People expect the item somewhere else. Move the item, or rename the section they gravitate toward. The first-click cluster tells you exactly where they expect it.
  • High success, low directness. The item is findable but the path is not obvious. Usually a labeling or nesting problem one level up. Tighten the parent label or flatten the hierarchy.
  • Split first clicks across two sections. Two categories overlap in users’ minds. Consider merging them, or making the distinction between them clearer.
  • Low success, scattered paths. No shared mental model exists for this item. This is the case for a fresh card sort to rebuild that branch from user input.

Make the changes, then retest the tasks that failed. A second tree test on a revised structure is fast and confirms whether your fixes actually moved the numbers. Treat tree testing as iterative rather than a one-time gate. When you are ready to move from a validated tree into a designed interface, follow up with moderated sessions, and lean on techniques from analyzing user interview data to make sense of the qualitative reactions design surfaces. For teams building a repeatable practice, our complete guide to user research and the B2B user research playbook place tree testing within a full research program.

Where tree testing fits in the bigger picture

Tree testing is deliberately narrow. It tells you whether your structure works, and nothing else. It will not tell you if the feature is valuable, if the visual design is clear, or if the copy persuades. That focus is a strength: by controlling for everything except structure, it gives you an unambiguous read on information architecture that no full-fidelity usability test can match, because in a real prototype you can never be sure whether a failure came from the navigation or the noise around it.

Sequence it well. Card sort to build, tree test to validate, usability test to confirm in context. Run the tree test before design so a bad structure never reaches a screen. And recruit participants who match your real audience, because a tree test with the wrong people is a precise measurement of the wrong thing. Methodology bodies such as the Nielsen Norman Group have documented the practice for years, and the pattern holds across products: the teams that measure findability early ship navigation that users never have to think about.

Ready to validate your information architecture with people who match your real users? Recruit verified participants on CleverX and get success rates you can trust.

Frequently asked questions

What is a tree test?

A tree test is an information architecture study where participants try to complete find-it tasks using only your navigation structure, stripped of visual design, search, and page content. You give them a task like where would you update your billing details, and they click through the text-only hierarchy until they land on an answer. It measures whether your labels and grouping let people find things, isolated from everything else on the page.

How is tree testing different from card sorting?

Card sorting is generative and asks participants to group and label content the way they expect, which helps you build a structure. Tree testing is evaluative and asks participants to find items in a structure you already have, which tells you whether it works. Card sort first to shape the tree, then tree test to validate it. Many teams run a card sort, draft the architecture, and confirm it with a tree test before development.

How many participants do I need for a tree test?

Plan on 30 to 50 participants per audience segment for reliable success rates. Because tree testing is quantitative, small samples produce noisy percentages that swing wildly when one person gets lost. If you are comparing two tree versions in an A/B setup, size each group to at least 30. For narrow B2B roles where the pool is small, 20 to 30 verified participants still gives a usable directional read.

What metrics does a tree test report?

The core metrics are success rate, the share of participants who reach a correct answer; directness, the share who got there without backtracking; and time to complete. Most tools also give you a first-click breakdown and a path visualization showing where people went. Read success and directness together, because a high success rate with low directness signals a confusing structure that people eventually brute-force their way through.

When should I run a tree test?

Run a tree test whenever you are validating or redesigning navigation, menu labels, or content grouping before build. Common triggers include a site or app redesign, a merger of two products, adding a major new content area, or analytics showing people cannot find a key feature. Because it uses a text-only tree, you can test before any design or engineering work exists, which makes it a cheap early checkpoint.

Can you run a tree test on B2B or specialized audiences?

Yes, and you should when your product serves a specific role. A tree test is only trustworthy if the people clicking through it think about your domain the way your real users do. Testing enterprise security navigation with general consumers produces misleading success rates. Recruit verified participants who match your target role, industry, and seniority so the find-it behavior reflects genuine mental models rather than guesses.