Mixed Methods: The Benefits of Combining Video and Audio in Quantitative Surveys

In this piece
Mixed methods research at its most useful generates a finding and its explanation in the same study. Adding video and audio responses to a quantitative survey is one of the most practical ways to get there. The rating scale shows where sentiment landed and the recorded response explains why.
Key Takeaways
- Video and audio responses carry tone, pacing and visual context that a text box cannot hold.
- Offering a recorded answer alongside text lets respondents who prefer speaking give fuller answers without losing those who would rather type.
- Speaking can feel more natural than typing, but a camera adds its own pressure, so let respondents choose their channel.
- Reading rating scales against recorded answers from the same respondents surfaces contradictions a single format would hide.
- Automated transcription and a researcher-reviewed first coding frame make audiovisual open-ends practical for large quantitative studies.
Reading Quantitative Scores Against Recorded Responses
In survey research, mixed methods means collecting rating scales and response counts alongside video or audio open-ends. The analyst then reads the two data streams against each other.
The value is corrective. Picture a respondent who rates a packaging redesign 5 out of 7 and says "yeah, it's fine," but pauses before answering and trails off at the end. The number files that answer as mildly positive. The words agree. Only the delivery says otherwise. In a text box, that answer reads as approval. In the recording, the hesitation is the finding.
This matters most in studies where the stakes are high enough that a misread finding costs real money. Concept tests, brand health tracking and post-launch diagnostic work all carry that risk. Running a purely quantitative wave and commissioning a separate qualitative phase six weeks later to explain a confusing number is slow. The two data sets also rarely align cleanly because the respondent pool, timing and stimulus have all shifted. Building video and audio response options into the quantitative instrument captures the explanation from the same respondents in the same fielding window.
The traditional quantitative survey asks the respondent to translate a reaction into a number or a typed sentence. A pause before answering or a voice that trails off carries meaning that a text field cannot hold.
Making Recorded Answers an Option
Open-ends are where surveys lose effort. Typing a considered paragraph on a phone is work. Many respondents keep it short. Some would rather speak for thirty seconds than type. A video or audio option gives them that route.
The research on voice answers is a useful reality check. Studies of smartphone surveys have found that a substantial share of respondents can't or won't record an answer, with item nonresponse well above what text open-ends see. A recorded answer therefore works best as an option alongside the text box. Respondents who want to speak get a fuller channel, while those who prefer typing lose nothing.
Being recorded also carries its own pressure. Speaking can feel more natural than typing for some respondents. A camera is more identifying than a text box, though, so the pull toward the acceptable answer does not disappear. Letting respondents choose audio over video or text over both keeps that pressure in check.
Reading Recorded Answers at Scale
Surveys that collect video, audio and text in parallel give analysts more than one way to check a finding. If respondents rate an experience highly but hesitate, qualify or contradict themselves when they talk about it, that tension is a finding in itself. Single-format studies tend to paper over exactly that kind of contradiction.
Some questions are hard to answer in a text box at all. Emotional experiences and perceptions tied to identity ask respondents to do interpretive work that a few typed words rarely capture. On sensitive subjects like health or money, audio often works where video does not. A text option should always remain available.
The practical barrier to audiovisual open-ends has historically been analysis time. Processing hours of recorded responses by hand is expensive, the same constraint that makes scaling qualitative research hard in any format. Enumerate transcribes each recorded answer at 99.5% accuracy and analyzes the visual content of video and images, such as the setting a response was recorded in. Sentiment comes from what respondents say, read from the transcript. The platform then proposes a first coding frame that the researcher reviews, renames or merges before it applies.
With transcription and a first coding frame in place, multimedia open-ends become practical for large quantitative studies where the richest comparisons live. In concept testing or brand health work, the payoff is a dataset where quantitative patterns can be checked against recorded evidence. Every theme stays linked to the response, transcript passage or clip behind it.
Want to see how video and audio responses work inside a quantitative study? Book a demo with Enumerate.
Related reading

Limitations of AI Content Analysis: What It Gets Right and Where It Breaks
A senior researcher's honest look at AI content analysis limitations: where it excels, where concept drift and contradiction handling go wrong, and what to watch for.
Read more
Descriptive Coding in Qualitative Research: A Practical Guide
Learn how to do descriptive coding in qualitative research, when to use it over thematic coding, and how to scale it without losing consistency.
Read more
Monadic vs Sequential Monadic Testing: Order Effects, Sample Economics, and Why Teams Default Wrong
Most teams pick the wrong product test design by default. Here's what the order effect research actually says about monadic vs sequential monadic testing.
Read more