The Right Sample Size for AI-Moderated Interviews

In this piece
Every proposal puts a number in the sample box, and if you are honest, it came from what the budget would carry as often as what the design needed. Twenty interviews. It has the ring of something settled in a journal somewhere. Mostly it is a budget constraint that hardened into a rule, and AI moderation is about to expose how much work it was doing.
Key Takeaways
- Saturation is segment-specific. What saturates one audience will not saturate five, and averaging across them is not a sample, it is an approximation.
- Code saturation and meaning saturation are different thresholds. Your themes show up early; the nuance behind them takes considerably longer.
- Probe quality sets your effective sample size. A moderator that accepts thin answers pushes meaning saturation further out, whatever the n.
What the Saturation Research Actually Found
The convention is a rounding of something more specific, worth knowing if you have to defend it to a client. Guest, Bunce, and Johnson's foundational study analyzed 60 interviews and found saturation arrived within the first twelve, with high-level themes visible as early as six.
It is also half the picture, and the missing half is the one that bites you in the debrief. Hennink, Kaiser, and Marconi drew a distinction the convention ignores. Code saturation, where no new themes appear, comes early. Meaning saturation, where no further dimensions or nuances emerge, takes considerably longer, in their work 16 to 24 interviews. Stop at the first threshold and you have a complete list of topics and nothing interesting to say about any of them, the study that reads fine on the summary slide and falls apart when someone asks why.
Saturation Is Segment-Specific
You already know where this breaks. A study labeled "national" with 20 interviews across five segments has four per segment. Four. And someone will still ask what lapsed users thought, and someone will still answer. That is sampling anxiety managed by averaging.
And this is measured, not asserted. Hagaman and Wutich revisited the Guest study across four cross-cultural sites and found 16 or fewer interviews identified common themes within a homogeneous group, while themes cutting across all sites required 20 to 40. Within a segment the small numbers hold. Across segments they do not.
What Scale Actually Buys
The weak argument for more AI-moderated interviews is that they are cheaper. The one that should interest you is that scale changes what you can claim in the readout.
Take a repositioning test across loyalists, lapsed users, and new-to-category considerers. Traditional pricing gets you eight per cohort, under the threshold for all three, and you write the deck anyway with careful hedging around the thin cells. At a fraction of the cost each cohort reaches saturation on its own terms, and the finding stops being an average across an implicit blend. The same economics reach the secondary markets that get dropped for cost reasons rarely stated in the methodology, and turn a second wave into a research decision rather than a budget exception. That is the real unlock in qual at scale: coverage of the audiences that used to get skipped.
The Constraint Moved From Budget to Design
The comfortable line, and the one most vendors give you, is that the saturation logic is unchanged and only the cost differs. That is not quite right.
Meaning saturation depends on how hard the moderator pushes. Nuance surfaces because someone chased a hedge or held an earlier claim against a later one. A moderator that accepts "it's fine, I guess" and moves on hits code saturation on schedule and meaning saturation late, or never. Probing quality is therefore a sample-size variable: weak probing does not just lower quality at a given n, it raises the n you need to learn anything. Evaluate AI-moderated interviews on probe depth and you are answering a sampling question, not a feature one.
Recruitment runs the same way in reverse, and this one should worry you more. Scale multiplies whatever is already in your sample. If the screener is loose or the panel is thick with professional respondents, 200 interviews will not rescue the study, they will hand you a more confident version of the same error. Solving a design problem with volume works even less well when volume is easy.
What the Research Implies for Study Design
None of this yields a number you can drop into a proposal. What it gives you is a set of questions worth answering before you set the sample.
- Within a segment, the evidence clusters low for themes and higher for nuance, meaning saturation landing in the high teens to mid twenties. Where you sit depends on how much of your value is the nuance.
- Across segments, the counts do not pool. The question is whether cross-cutting themes are in the brief, because that is where the requirement climbed to 20 to 40.
- For longitudinal work, establishing themes is a different task from detecting whether they moved, and the second has tolerated lighter waves.
The bottleneck was never the interview count. It was the cost of reaching the right count in every segment that mattered, and that has changed. What has not is the part that was always yours: the count only pays off if the probing is deep enough to earn it and the recruitment clean enough to trust it. The budget excuse is gone. The design judgment is not.
Curious how this applies to your study design? Book a demo with Enumerate.
Related reading

Diary Study Best Practices: Stop Optimizing the Wrong End
Most teams invest in AI analysis but lose diary studies at recruitment. Here's what actually determines whether a diary study succeeds or fails.
Read more
Panel Fraud and AI Moderated Research: What's Actually Changed
Panel fraud predates AI moderation and survives it. Here's what the fraud taxonomy looks like, what AI makes worse, and what real detection architecture requires.
Read more
AI Based Content Analysis for Qualitative Interviews
Learn when and how to use AI based content analysis for qualitative interviews — practical guidance on process, fit, and where AI genuinely changes the workflow.
Read more