29 August 2026
A user interview script for customer research is not a list of questions to read aloud — it's a spine of five or six fixed prompts plus a set of banned questions, built so twenty different conversations still produce comparable data. The script survives contact with a real participant when it protects the behavior you're after even when the person tries to hand you a feature request instead.
The first ninety seconds set the contract for the whole interview: what you're asking permission to record, why, and what happens to it. Skip this and half your participants will hedge their real opinions because they don't know who's going to read the transcript.
Say the same four things every time, in the same order, so participants across your twenty interviews all start from the same footing: what the session is for, that you're recording, who sees the recording, and that there are no wrong answers because you're not testing them.
Some questions feel productive and return nothing but noise, because they ask the participant to predict their own future behavior or price something that doesn't exist yet. "Would you use this?" gets a polite yes almost every time — agreeableness, not intent.
Ban future-tense and hypothetical-price questions from the script entirely. Replace them with a request for the last time they actually did the thing, which is the only data point that predicts what they'll do next time.
Somewhere around interview six, someone will stop answering and start pitching: "you should just add a button that..." This is the participant trying to be helpful, and it's the moment a script without a plan falls apart, because the interviewer either argues or writes down a feature request as if it were data.
The fix is one sentence, said the same way every time: "That's a good idea — before we get there, tell me about the last time the current version let you down." It moves the conversation back to a specific past event, which is the only thing you can actually code and compare across interviews. Their solution isn't wrong, it's just not evidence; the problem underneath it is.
Internal usability tests run under whatever policy your company already has. Customer interviews with people outside the company don't — you need spoken consent every time, and in two-party consent jurisdictions (most of the US West Coast, parts of the EU under stricter interpretations of GDPR) implied consent from showing up on the call is not enough.
Get the consent on the recording itself, out loud, before the first real question: "Just confirming out loud, it's okay if I record this for our internal research." A platform that transcribes the call itself, the way AVAY does, means there's no separate bot joining to explain — the participant just needs to hear the one sentence and answer it. One limitation worth planning around: if the session is also recorded as video, that file saves to the machine that started the meeting, not a shared cloud store, so build your storage and redaction step around that before the interviews start, not after.
A summary written during the interview is already an interpretation — "she said the export feature was confusing" throws away whether she said "confusing," "buggy," or "I gave up and emailed the file to myself instead," which are three different problems with three different fixes.
This is the actual argument for verbatim transcripts over live notes: you can't code language you didn't keep. Searching a transcript for the phrase a participant actually used, three interviews later, is how you notice that four different people independently said "gave up and emailed it to myself" — a pattern a paraphrased note would have flattened into "found it hard to use."
Twenty interviews are comparable when they share a spine — the same five or six questions asked in the same order, every time — and different when each interviewer follows their own curiosity into the gaps. Both are necessary. All spine and no follow-up gives you shallow, repetitive data; all follow-up and no spine gives you twenty different studies you can't line up next to each other.
Write the spine questions word-for-word in the script and require interviewers to ask them exactly as written, out loud, even if it feels stiff. Everything after each spine question — the probes, the "tell me more," the tangents you chase — stays unscripted. That's the seam where comparability and depth actually coexist.
| Comparability across interviews | Risk of leading the answer | Depth per session | |
|---|---|---|---|
| Fully scripted (read every question verbatim, no deviation) | High | Low | Low — no room to follow a thread |
| Fully open (conversation, no fixed questions) | Low | Medium — interviewer curiosity steers it | High |
| Spine + free probes (fixed core, open follow-up) | High on the spine questions | Low on the spine, managed on probes | High |
Most teams see the same three or four themes repeat by interview eight to twelve, with fewer new themes appearing after that. Fewer than five interviews usually isn't enough to distinguish a real pattern from one talkative participant.
The spine questions, yes, word for word — that's what makes twenty interviews comparable instead of twenty separate stories. The follow-up questions should differ by interviewer, because they're chasing what that specific participant said.
Send the topic, not the questions. If they prepare answers to the exact wording, you get a rehearsed version of their opinion instead of a description of what they actually did.
Interview them anyway with notes only, and mark that transcript as note-based rather than verbatim in your analysis so you don't quote it as if it were exact language. Losing the recording is better than losing the participant.
Keep the spine identical if you want the two rounds to compare cleanly, and change only the probes. If you rewrite the spine questions between rounds, you're running two different studies, not one study twice.
A script that survives a real participant has a fixed spine you never reword, a short list of questions you never ask, and a plan for the moment someone starts pitching features instead of describing behavior.
AVAY is a video meeting platform that transcribes the call itself — no bot joins, because there is nothing to join. Start one at avay.ai, read how each part works in the documentation, or see what it costs.