How to run a MaxDiff survey and read the results
MaxDiff forces respondents to choose the best and worst item in each set, giving you a cleaner priority ranking than a rating scale. Here is how to design, field, and score one.
MaxDiff is a survey method that shows respondents small sets of items and asks them to pick the best and worst in each set. Repeat that across enough sets and you get a clean, scaled ranking of every item on your list. It is the fastest reliable way to answer a question every product and marketing team faces: of all these features, benefits, or messages, which ones actually matter most to the people we are building for?
This guide walks through what MaxDiff is, when to use it instead of a rating scale or conjoint, how to design the item list and choice sets, how many responses you need, and how to score and read the results so they drive a decision.
What a MaxDiff survey is
MaxDiff, also called best-worst scaling, was developed by Jordan Louviere in the 1990s. Instead of asking people to rate each item on a scale, you show them a subset of items, usually four or five at a time, and ask two questions: which is best (most important, most appealing) and which is worst. Each respondent sees many of these sets, with the items rotating so every item is evaluated several times against different competitors.
The core idea is trade-off. When you ask someone to rate ten features on a 1-to-5 scale, most people rate almost everything a 4 or 5, because in isolation everything sounds nice. That gives you scores clustered at the top with no real separation. MaxDiff removes that escape hatch. In every set, one item has to be best and one has to be worst, so respondents are forced to reveal their true priorities.
The output is an interval-scaled ranking. You not only learn that feature A beats feature B, you learn by roughly how much, which is what lets you build a defensible priority list.
MaxDiff vs a rating scale vs conjoint
Teams often reach for a rating scale out of habit, or a conjoint study because it sounds rigorous. Here is how the three compare so you can pick the right tool.
| Dimension | Rating scale | MaxDiff | Conjoint analysis |
|---|---|---|---|
| What it measures | Absolute importance of each item | Relative priority across items | How attributes and price combine into choice |
| Respondent task | Rate each item independently | Pick best and worst in each set | Choose between product bundles |
| Discrimination | Low; scores cluster at the top | High; forces clear trade-offs | High; models realistic choices |
| Best for | Quick pulse checks | Ranking features, messages, benefits | Pricing, packaging, product configuration |
| Complexity to design | Low | Medium | High |
| Typical sample size | 100 or more | 200 to 300 | 300 or more |
| Scale bias risk | High; varies by person and culture | Low; ranking is comparative | Low |
Use a rating scale when you need a fast temperature check and separation does not matter. Use MaxDiff when you have a list of standalone items and you need a trustworthy priority order. Use conjoint when items trade off against each other and against price, and you need to simulate what people would actually buy. If you are still mapping which method fits your question, our complete walkthrough to product research methods, frameworks, and best practices lays out the full landscape.
MaxDiff also pairs well with prioritization frameworks. The Kano model feature prioritization guide helps you classify features by how they affect satisfaction, while MaxDiff tells you the relative weight of each one. Running both gives you a two-dimensional view: what type of feature it is, and how much people want it.
Step 1: Write a clean item list
Everything downstream depends on the quality of your item list. A few rules keep it usable.
Keep items comparable. Every item should sit at the same level of abstraction. Do not mix a broad benefit like “saves me time” with a narrow feature like “keyboard shortcut for exporting.” Respondents cannot trade those off meaningfully.
Keep items independent. Items should not overlap or imply each other. If two items say nearly the same thing, they split votes and both look weaker than they are.
Use consistent phrasing. Start each item the same way and keep length similar. Long, detailed items can win simply because they sound more complete, which is a wording artifact, not a real preference.
Aim for 10 to 30 items. Fewer than 10 and you probably do not need MaxDiff. More than 30 and the survey gets long, though modern designs can handle larger lists by showing each respondent only a subset. For most feature or message tests, 15 to 25 items is the sweet spot.
If your list is a jumble of features, messages, and outcomes, split it into separate studies. A MaxDiff on feature priorities and a separate MaxDiff on message appeal will each produce cleaner results than one mixed study. For message work specifically, combine MaxDiff with qualitative validation using our guide on how to run message testing with real buyers.
Step 2: Design the choice sets
Once you have your items, you build the experimental design that decides which items appear together in each set. You do not have to do this by hand; MaxDiff tools generate a balanced design automatically. But you should understand the parameters you are setting.
Items per set. Four or five is standard. Fewer than four gives you too little information per screen; more than five makes the trade-off harder and slows people down.
Sets per respondent. Usually 10 to 20. The guiding principle is that each item should appear at least three times per respondent, ideally more, so the model can estimate a stable score for it. A quick check: multiply items per set by number of sets, then divide by total items to get average appearances per item.
Balance and orthogonality. A good design shows each item roughly the same number of times (balance) and pairs items together evenly so no two items are always or never seen together (orthogonality). Balanced designs prevent artificial inflation or deflation of any single item.
For a 20-item list, a typical design might show 4 items per set across 15 sets, giving each item three appearances per respondent. That keeps the survey to a few minutes while producing reliable scores.
Step 3: Decide sample size
Sample size in MaxDiff is less about hitting a magic number and more about who is answering.
For a single overall ranking, 200 to 300 qualified respondents is a solid target. Because MaxDiff aggregates many judgments per person, it extracts a lot of signal from each respondent, so you need fewer people than an attitudinal survey might.
If you plan to compare segments, size for the segments, not the total. Aim for roughly 150 to 200 respondents per segment you want to read separately, whether that is enterprise versus SMB, or one persona versus another. Undersized segments produce noisy rankings that flip on replication. If segmentation is central to your study, plan it up front using our guide on how to run a customer segmentation study.
The bigger reliability driver is audience fit. A MaxDiff fielded to 500 random panelists who do not use your category will give you a precise ranking of what irrelevant people think. A MaxDiff fielded to 200 verified buyers in your exact segment will give you a ranking you can act on. This is where recruitment quality matters more than raw count, and it is worth reading how to recruit B2B research participants before you field, especially for niche or senior audiences.
Step 4: Field the survey to the right audience
MaxDiff results are only as good as the people answering. This is the step teams most often shortcut, and it is where studies quietly go wrong.
CleverX is a B2B research platform with more than 8 million verified professionals plus B2C reach across 150 or more countries. Every participant is identity- and employment-verified, so when you run a MaxDiff on developer tooling features, you are hearing from real developers, and when you test procurement messaging, you are hearing from real procurement leaders. Studies run pay-as-you-go with roughly 2 to 5 day delivery, so you can field a feature-priority MaxDiff early in a sprint and have scored results before planning locks.
The reason this matters for MaxDiff specifically: the method needs enough responses from your true target audience to be reliable. A ranking built on the wrong sample looks just as clean and confident as a ranking built on the right one, which makes bad targeting dangerous. Verified recruitment removes that risk. For a broader grounding in fielding practices, the basics of market research is a useful primer for anyone new to sampling and screening.
Keep the survey itself short. MaxDiff is repetitive by design, so 10 to 15 sets is comfortable, but push past 20 and quality drops as people speed through. Add a couple of attention or consistency checks, and drop respondents who contradict themselves across repeated sets.
Step 5: Score the results
There are two levels of scoring, from simple to rigorous.
Best-minus-worst counting
The quickest score is counting. For each item, count how many times it was chosen best across all respondents, count how many times it was chosen worst, and subtract worst from best. Rank items by that net score. Items with a high positive score are clear priorities; items near zero are contested; items with negative scores are consistently rejected.
Best-minus-worst is easy to compute and easy to explain to stakeholders, which makes it a great first cut. Its limit is that it is a count, not a probability, so the gaps between items are harder to interpret precisely.
Utility scores with hierarchical Bayes
For a probability-based view, most MaxDiff tools estimate utility scores using hierarchical Bayes (HB). HB models each respondent individually while borrowing strength from the whole sample, producing a utility for every item for every person. You can then average these and rescale so all items sum to 100.
Rescaled utilities are the most useful output for decisions. If feature A scores 18 and feature B scores 6, you can say feature A is roughly three times as important to this audience, which is a far stronger statement than “A ranked higher than B.” Sawtooth Software, which popularized commercial MaxDiff, publishes a clear overview of MaxDiff analysis and hierarchical Bayes if you want the statistical detail.
Whichever method you use, look at the spread, not just the order. Sometimes the top five items are far ahead of everything else and the choice is obvious. Sometimes items two through eight are bunched together, which tells you the priority is genuinely contested and you may need qualitative follow-up to break the tie.
Step 6: Turn scores into decisions
A ranked list is not a decision. To make MaxDiff pay off, connect it to what your team will actually do.
Draw the line. Look for the natural break in the scores and treat items above it as the priority set for the next cycle. If the gap between rank 5 and rank 6 is large, that is a clean cut point.
Read segments before you commit. An overall ranking can hide a split where enterprise buyers want governance features and SMB users want speed. If your segments disagree sharply, a single roadmap may not serve both, and that is a strategic finding in itself.
Pair rankings with reasons. MaxDiff tells you what people prioritize, not why. Run a handful of interviews on the top and bottom items to understand the motivation, so you build the right version of the winning feature.
For moving from a scored list to a shipped decision, our guide on how to turn product research into better product decisions covers how to socialize findings and avoid the trap of research that never changes anything. And if you want to place MaxDiff inside a broader, repeatable process, the market research methodology guide shows how method selection, sampling, and analysis fit together across a research program.
Common mistakes to avoid
- Mixing item types. Features, benefits, and messages in one list produce a ranking no one can act on. Split them.
- Overlapping items. Near-duplicate items split votes and both look weaker. Edit for independence before fielding.
- Too many items, too few sets. If items appear fewer than three times per respondent, scores get noisy. Balance the design.
- Wrong audience. The single biggest error. Precise scores from the wrong people are worse than useless, because they look trustworthy.
- Reading order without spread. Always check the size of the gaps between items, not just their rank.
Bring MaxDiff to the right people
MaxDiff is one of the most reliable ways to turn a long wishlist into a defensible priority order, but it depends entirely on hearing from the audience you actually build for. When you are ready to field, you can recruit verified participants on CleverX and reach the exact professionals or consumers your product serves, with results back in days rather than weeks.
Frequently asked questions
What is a MaxDiff survey?
MaxDiff, also called best-worst scaling, is a survey method where respondents see small sets of items and pick the best and worst in each set. Repeating this across many sets produces a ranked, scaled priority order for the full item list without the flattening you get from rating scales.
How is MaxDiff different from a rating scale?
A rating scale lets people mark everything as important, so scores cluster near the top and give you little separation. MaxDiff forces a trade-off in every set, so respondents must decide what matters most and least. The result is clearer discrimination between items and rankings that hold up across cultures.
How many items and choice sets should a MaxDiff survey have?
Most MaxDiff studies test between 10 and 30 items, show 4 to 5 items per set, and ask each respondent to complete 10 to 20 sets. A common rule is that each item should appear at least 3 times per respondent so the model can estimate a stable score for it.
How many respondents do you need for MaxDiff?
For a whole-sample ranking, 200 to 300 qualified respondents is a common target. If you plan to compare segments, aim for roughly 150 to 200 per segment. The bigger driver of reliability is that responses come from your real target audience, not from an unscreened panel.
How do you score MaxDiff results?
The simplest score is best-minus-worst: count how often an item was chosen best, subtract how often it was chosen worst, and rank items by the result. For a probability-based view, use hierarchical Bayes to estimate utility scores that can be rescaled to sum to 100 across all items.
When should you use conjoint instead of MaxDiff?
Use MaxDiff to rank standalone items such as features, messages, or benefits when they do not trade off against price or each other. Use conjoint when you need to model how bundles of attributes and price levels combine into a product choice, since conjoint simulates realistic buying decisions.