How many user interviews do you actually need?
Every research plan meeting eventually reaches the same question: how many interviews are enough? The most common answer — five — is quoted so often that it has detached from what it actually means. So let's start there, and then get to the honest answer, which is: it depends on your segments and your stakes, and the constraint that produced the old rules of thumb is disappearing.
Where the 5-user heuristic comes from
The famous number comes from usability research by Jakob Nielsen and Tom Landauer in the early 1990s. Their finding, roughly: in a usability test of a single interface, the first five users will surface around 85% of the usability problems, and each additional user surfaces less that is new. Their recommendation was not "five is enough research" — it was that with a fixed budget you learn more from three rounds of five-user tests, iterating between rounds, than from one fifteen-user test.
Note what the heuristic assumes: one homogeneous user group, one interface, and a goal of finding usability problems — defects that most users will trip over. Under those conditions, small samples genuinely work, because you're fishing in a pond where the fish are everywhere.
What the heuristic doesn't cover
Most interview research isn't usability testing. If you're doing customer discovery, win/loss analysis, market understanding, or diligence, you're not looking for defects that everyone hits — you're mapping a distribution of experiences, motivations, and contexts. And distributions need coverage.
Three things move the number up:
- Segments. The five-user math applies per segment. If enterprise buyers, mid-market buyers, and end users have genuinely different contexts, five interviews total gives you fragments of three stories rather than one complete story. Five per segment is the honest floor — and now you're at fifteen or twenty.
- Stakes. Deciding the color of a button and deciding whether to enter a market can both be informed by interviews, but they don't deserve the same evidence standard. When a decision is expensive to reverse, the cost of a wrong conclusion dwarfs the cost of ten more conversations.
- Variance. Some questions have convergent answers — ask ten accountants how they close the books and you'll hear the same workflow by interview four. Others are wildly divergent — ask ten founders why they chose their bank. The more varied the answers, the more interviews it takes to see the shape of the variation.
Saturation: the real stopping rule
The methodologically honest stopping rule isn't a number — it's saturation: you stop when new interviews stop teaching you new things. In practice, researchers operationalize it as something like "no new themes in the last three to five interviews." Empirical studies of thematic saturation in reasonably homogeneous groups often land in the 9–17 interview range, which is why experienced researchers wince at both "five is plenty" and "we need fifty."
Saturation has a catch, though: you can only detect it in hindsight, and only per segment. If you never interviewed the segment, you can't saturate it. A study can be fully saturated on the customers you talked to and silent about the ones you didn't.
What changes when interviews are cheap to parallelize
Here's the part that's genuinely new. Every heuristic above was shaped by one constraint: interviews were expensive. Each one cost a researcher-hour plus scheduling overhead, so the discipline developed rules for rationing — how few interviews can we get away with?
When AI voice agents run interviews in parallel, the marginal interview stops costing calendar time. Thirty 15-minute interviews complete in roughly the time of one. That doesn't mean "more is always better" — synthesis attention is still finite, and a sloppy guide asked 100 times is just a sloppy study at scale. But it changes the planning question in two useful ways:
- Stop rationing, start segmenting. Instead of asking "can we get away with five?", ask "which segments deserve their own saturation?" The budget that used to buy five total conversations can now buy proper coverage of each segment that matters — including the ones you'd previously have skipped and silently guessed about.
- Saturation becomes observable instead of hoped-for. With sequential interviews, teams stop when the calendar runs out and call it saturation. With parallel interviews, you can actually run past the expected saturation point cheaply, and see whether themes really did stop appearing — or whether interview 24 surfaced the objection that would have sunk the launch.
A short, honest answer
Usability test of one interface, one user type: ~5 per iteration round. Discovery or evaluative research within one segment: expect saturation somewhere around 10–15. Multi-segment or high-stakes decisions: aim for saturation per segment, not in aggregate — and when interviews are parallel and cheap, cover the segments you would once have rationed away.
The number was never really the point. Coverage of the people your decision touches is the point — the old rules of thumb were just the best coverage that sequential, human-scheduled interviews could afford.