September 9, 2026 · By The Listently Team
How to Transcribe Qualitative Interviews for Coding

To transcribe qualitative research interviews for coding, record clean audio (one recorder, minimal crosstalk, a stated verbal consent preamble), upload the file to an automatic transcription tool that does speaker diarization, rename the detected speakers to your study’s participant IDs (P07, Interviewer, etc.) rather than real names, correct the transcript against the audio while listening, then export to Word. The Word file is the format nearly every coding workflow accepts: NVivo, ATLAS.ti, MAXQDA, and Dedoose all import .docx, and manual coding in margins or a spreadsheet starts from the same file. The two decisions that matter most happen before coding: how much verbatim detail your analysis method actually needs (discourse analysis needs pauses and fillers; thematic analysis usually does not), and whether your speaker labels are consistent enough that a quote can be traced back to a participant six months later. Get those right and transcription is a mechanical step. Get them wrong and you re-do the corpus. Budget roughly three to five hours of correction time per hour of audio if you’re transcribing from scratch, or 30 to 60 minutes per hour of audio if you’re correcting a machine transcript.
Decide your transcription level before you touch the audio
Not all qualitative transcription is the same, and the level you choose determines how long correction takes.
- Verbatim (clean/intelligent verbatim): every word spoken, minus false starts, stutters, and most “um”s. Standard for thematic analysis, framework analysis, grounded theory, and most market research. This is what you want unless you have a reason otherwise.
- True verbatim: includes fillers, repetitions, laughter, and audible non-speech. Used when hesitation itself is data — sensitive topics, some IPA work.
- Jefferson or similar notation: timed pauses, overlap markers, intonation, stretched sounds. Required for conversation analysis and discourse analysis. No automatic tool produces this; you build it on top of a base transcript.
Pick one level and write it into your methods section before the first interview, because mixing levels across a corpus makes cross-case comparison unreliable. Automatic transcription gets you to clean verbatim quickly; anything above that is manual work layered on a machine draft.
Recording quality determines how much correcting you’ll do
Machine transcription accuracy on a quiet one-on-one interview with a decent microphone is dramatically better than on a phone speaker recording in a café. The difference shows up as hours of your time.
Practical minimums for research interviews:
- Record in a room with soft surfaces and no HVAC hum directly overhead. Carpet and curtains help more than expensive equipment.
- Use one recorder placed between speakers rather than each person recording separately, unless you plan to sync tracks.
- For remote interviews, ask participants to use wired earbuds with a mic — laptop speakers cause echo that diarization struggles to separate.
- Ask people not to talk over each other, and gently re-ask questions when they do. Overlapping speech is the single biggest cause of mislabeled speakers.
- Have each participant say their study ID at the top of the recording. It makes matching files to consent forms trivial later.
We wrote a longer version of this list in 9 Ways to Improve Transcription Accuracy Before You Record, and the related rules for group settings are in Multi-Speaker Transcription: 7 Rules for Accuracy.
Thirty seconds of setup before recording saves an hour of correction per interview, which across a 25-participant study is most of a working week.
Consent and ethics come before the recorder, not after
Recording-consent rules vary by jurisdiction — some places require only one party’s consent, others require all parties to agree, and rules differ again for phone and video calls. On top of that, your IRB, ethics committee, or client contract will have its own requirements that are usually stricter than the law.
For qualitative research specifically, check three things before you record:
- Where the audio and transcripts are stored and processed, and whether that satisfies your data management plan. Many ethics protocols require you to name the transcription service.
- Whether the vendor uses your data for anything beyond producing your transcript. Listently doesn’t use uploaded audio or transcripts to train any AI model, doesn’t share them with third parties, and deletes the uploaded audio immediately after processing — we go into why that matters in Does Transcription Software Train AI on Your Audio?.
- Your de-identification plan, including who holds the linking file between real names and participant IDs, and where it lives (not in the same folder as the transcripts).
Record a short verbal consent preamble at the start of every interview even when you have a signed form — it timestamps consent inside the same artifact as the data.
Upload, then rename speakers to participant IDs before anything else
Automatic diarization labels speakers generically: Speaker 1, Speaker 2. Those labels are useless for analysis and dangerous for citation, because after ten interviews you can’t tell which Speaker 2 is which.
The fix is to rename them immediately after the transcript comes back and before you start correcting. In Listently, speaker labels are editable, and renaming a speaker applies everywhere the label appears — including inside the AI summary — so you rename once rather than find-and-replacing through a long document.
A naming convention that survives contact with a real study:
| Role | Label to use | Avoid |
|---|---|---|
| Interviewer | INT or Interviewer |
Your initials, if multiple people interviewed |
| Participant | P07 (zero-padded, matches your master log) |
Participant, first names, P7 |
| Second participant in a dyad | P07a / P07b |
Speaker 2 |
| Translator or interpreter | INTERP |
Merging them into the participant label |
Zero-pad your participant numbers (P07, not P7) so files and quotes sort correctly, and never type a real name into the transcript at all — the transcript should be de-identified from the moment it exists.
If a participant says their own name or a colleague’s name mid-interview, replace it in the transcript text with a bracketed placeholder like [COLLEAGUE NAME] during correction. Doing this at the transcript stage rather than at the write-up stage means you never have an identifiable copy sitting in your project folder.
Teach the tool your jargon so you’re not fixing the same word 200 times
Qualitative corpora are full of terms a general model will mangle: drug names, program acronyms, regional place names, brand names in market research, theoretical vocabulary. The same misspelling then appears in every transcript, and every coded quote inherits it.
Listently’s custom vocabulary lets you save names, jargon, and acronyms so future transcripts spell them correctly automatically — 15 terms on the Free plan and 50 on Pro. Load it before you transcribe the corpus, not after.
Worth saving:
- Your study’s program or intervention name
- Recurring acronyms (a health study might need
PROM,MDT,ICS) - Product and competitor brand names in market research
- Place names and organization names participants mention repeatedly
- Any specialist term your discipline uses that a general audience wouldn’t
Setup details are in How to Set Up Custom Vocabulary: Fix Names & Jargon. Spending ten minutes on a vocabulary list before transcribing 20 interviews is the highest-leverage step in this whole workflow.
Correct the transcript while listening — that pass is also your first analysis
Never ship a machine transcript straight into coding software. Play the audio at 1.0x to 1.25x and read along, fixing errors as they appear. Most researchers find this pass doubles as immersion: you start noticing themes, and it’s worth keeping a separate memo file open while you do it.
What to check specifically:
- Speaker attribution at turn boundaries. Diarization errors cluster where one person interrupts another; a misattributed sentence becomes a misattributed quote.
- Numbers, dates, and quantities, which machines get wrong more often than words.
- Negations. “I didn’t feel supported” versus “I did feel supported” is a coding-level error.
- Anything you can’t hear clearly — mark it
[inaudible 00:14:22]with the timestamp rather than guessing. - Names of people and places, replaced with bracketed placeholders per your de-identification plan.
Do not silently clean up grammar into standard English — for many research questions, how participants phrase things is part of the data.
Export to Word, because that’s what coding software imports
Once corrected, export the transcript to Word. Listently exports to Word, PDF, and SRT/VTT subtitles; for coding workflows you want the .docx.
Why Word specifically:
- NVivo, ATLAS.ti, MAXQDA, and Dedoose all import .docx and preserve paragraph structure, which is what they use to segment text into codeable units.
- Several packages can auto-code by speaker if your labels are consistent paragraph-leading text — which is exactly why the renaming step earlier matters.
- If you’re coding manually, Word’s comment margin works as a codebook column, and the file prints cleanly for paper-based coding.
- PDF is the wrong choice here: it’s fine for archiving a final transcript with an appendix, but coding software handles it worse and you can’t easily edit it.
Before importing a whole corpus, run one transcript through your software end to end and check that speaker turns segment the way you expect. Fixing a formatting convention across 30 files is much cheaper than re-importing 30 coded files.
Practical workflow for a full study
For a 20-to-40 interview project, the sequence that wastes the least time:
- Build your custom vocabulary list from the interview guide and pilot interviews.
- Record interviews with verbal consent preambles and participant IDs stated aloud.
- Upload audio in batches. On Pro, you can select or drop up to 10 files at once instead of uploading one at a time, and you can paste a Google Drive share link to fetch a file directly, which matters if your recordings already live in a shared research drive.
- Rename speakers to participant IDs immediately on each transcript.
- Correct against the audio; memo as you go.
- Export each transcript to Word using a consistent filename pattern (
STUDY_P07_2026-03-14.docx). - Import to your coding software and verify segmentation on the first file.
One note on plan limits: the Free plan allows 3 transcriptions per day at up to 40 minutes per file, which is workable for a small pilot but not for a corpus of hour-long interviews. Pro removes the per-file length limit and the daily cap — the detail is in Listently Free vs Pro: When You’ll Hit the Wall.
If you want to test the workflow on a single pilot interview before committing, start with a free transcript and check the speaker labeling on your worst-quality recording, not your best one.
Common questions
Can I use automatic transcription in an IRB-approved study? Usually yes, but you generally have to name the service in your protocol and data management plan, and describe how audio is stored, processed, and deleted. Check with your committee before you record rather than after.
How long does transcription actually take per interview? Manual transcription from scratch typically runs three to five hours per hour of audio. Correcting a machine transcript of a clean one-on-one recording usually takes 30 to 60 minutes per hour of audio, and less once your custom vocabulary is dialed in.
Should I keep filler words like “um” and “you know”? Depends on your method. Thematic and framework analysis generally use clean verbatim without fillers. Conversation analysis, discourse analysis, and some IPA work need them, plus pause timings. Decide once and apply it to the whole corpus. If your method calls for clean verbatim, Listently Pro’s “Hide filler words” toggle strips them automatically without touching the underlying transcript, so you can switch back to true verbatim any time.
What if the recording has three or more speakers, like a focus group? Diarization gets harder as speakers increase and overlap grows. Ask participants to avoid talking over each other, seat the recorder centrally, and expect a longer correction pass on speaker attribution specifically. The general approach is covered in How to Transcribe an Interview: A Practical Guide.