August 19, 2026 · By The Listently Team
How to Transcribe an Interview: A Practical Guide
To transcribe an interview: record with a dedicated mic close to each speaker in a quiet room, save the file as WAV or MP3, upload it to an automatic transcription tool, let it detect and separate the speakers, rename the generic speaker labels to real names, then read through the transcript once to fix proper nouns and technical terms. Export the result as Word or PDF for publication, or SRT/VTT if the interview is going out as video. For a 45-minute interview, the machine part takes minutes and the cleanup pass takes roughly 10 to 20 minutes.
The quality of your transcript is decided before you ever upload anything. Everything below is ordered the way the work actually happens.
Step 1: Set up your recording for clean audio
Before you hit record: recording-consent laws vary by location — some places require only one party to consent, others require everyone on the call to. If you’re interviewing someone, it’s worth a quick check on the rule where you and your subject are, and telling them you’re recording either way.
Transcription accuracy tracks almost linearly with recording quality. A cheap mic six inches from someone’s mouth beats an expensive mic across a table.
Use a separate mic per speaker when you can. Two lavalier mics, two USB mics, or two phones recording locally will all outperform a single device in the middle of the table. Separated audio also makes speaker detection more reliable, because the algorithm has cleaner voice characteristics to work with.
Kill the room noise you can control. Turn off HVAC, close windows facing traffic, move away from refrigerators and fans. Background hum is the single most common cause of dropped words. Soft surfaces help: a room with a rug and curtains is better than a glass-walled conference room.
Watch for the things people forget. Phone notifications, a laptop fan spinning up, jewelry tapping the table, and paper shuffling all land in the recording. Ask your subject to silence their phone before you start, not after.
Record a 20-second test and actually listen back. Play it through headphones. If you can hear a hiss or the voices sound distant, fix it now — you cannot fix it later.
Choose a sensible format. WAV and FLAC are uncompressed and preserve the most detail. MP3 and M4A are compressed but perfectly adequate at 128 kbps or higher. If you’re recording video, MP4 or MOV is fine — a good transcription tool will pull the audio track directly, so you don’t need to convert anything first. Listently accepts MP3, WAV, M4A, FLAC, OGG, MP4, MOV, and MKV.
For remote interviews, record locally on both ends. Video call recordings compress audio aggressively and drop packets when connections wobble. Having each person record their own side on their own machine gives you two clean tracks. If that’s not practical, the platform recording will still work — just expect slightly lower accuracy.
Step 2: Upload and process the file
Once you have the file, upload it and let the tool do the first pass. This is the fast part. A typical interview comes back as a speaker-labeled transcript in minutes rather than the hours a manual pass takes.
If you’re working across languages, check that your tool handles the one you need. Listently supports over 100 languages for transcription, which covers most interview work outside of very rare dialects.
Two things worth confirming before you upload anything sensitive: whether your audio is used to train AI models, and how long files are retained. With Listently, audio and transcripts are never used to train any AI model and never shared with third parties, and uploaded audio is deleted immediately after processing. If you interview sources under confidentiality agreements, that distinction matters more than any accuracy percentage.
Step 3: Handle multiple speakers
Automatic speaker detection segments the audio by voice and assigns generic labels — Speaker 1, Speaker 2, and so on. Your job is to map those labels to real people.
Find each speaker’s first line. Skim the top of the transcript. In most interviews the interviewer speaks first, so Speaker 1 is usually you. Confirm by reading one or two of their lines — the person asking questions is easy to spot.
Rename the labels rather than find-and-replacing later. In Listently you rename a speaker once and it applies everywhere, including the AI summary. That’s much less error-prone than editing every instance by hand, and it means the summary reads as “Dr. Reyes noted…” instead of “Speaker 2 noted…”.
Expect some crosstalk errors. When two people talk over each other, any system will occasionally attribute a few words to the wrong person. These are usually short interjections — “right,” “mm-hmm,” “exactly.” Fix the ones that change meaning and ignore the rest.
Panel interviews need more care. With four or more voices, similar-sounding speakers get merged more often. If you’re recording a panel, ask each participant to state their name at the start. That gives you an anchor point in the transcript and makes the mapping obvious.
Step 4: Clean up the transcript
Do this as one focused pass with the audio playing alongside. Don’t try to perfect it in three separate rounds.
Fix proper nouns first. Names, companies, product names, and place names are the most common errors and the most damaging ones in a published piece. Search for each name you know should appear and correct every instance.
Then fix jargon and acronyms. Domain-specific vocabulary is the second most common error category. If your interview is about a niche subject, budget extra time here — or avoid the rework entirely: Listently lets you save a personal vocabulary list (names, product names, technical terms) that’s applied automatically on every future transcription, so recurring jargon comes back spelled correctly the first time.
Decide your cleanup level and stick to it. Verbatim transcription keeps every “um,” false start, and repetition — required for legal, research, and linguistic work. Clean verbatim removes filler words and stammers while keeping the actual words and structure — the right default for journalism and content. Edited transcription tidies grammar into readable prose — fine for internal notes, risky for direct quotation.
Timestamp anything you plan to cite. If you’ll need to point an editor or a fact-checker back to the audio, note the timecode next to key quotes as you go.
Use the summary as a checklist, not a replacement. Listently generates its AI summary from the actual transcript text and attributes points to specific speakers, and it’s instructed not to invent action items or conclusions that aren’t in the transcript. That makes it a useful map of the interview — scan it to find the sections worth reading closely first. It doesn’t replace reading the transcript before you publish a quote.
Transcript text and speaker names both stay editable after transcription completes, so you can correct things as you notice them rather than needing everything right in one sitting.
Step 5: Export for publication
Match the export format to the destination.
- Word (.docx) for anything going to an editor, a co-author, or a review process where tracked changes matter.
- PDF for archival copies, client deliverables, and anything you want to stay fixed.
- SRT or VTT for subtitles on a video version of the interview. Both are plain-text caption formats that upload directly to YouTube, Vimeo, and most video platforms.
If you’re publishing the interview as a Q&A article, export to Word and do your final editing there — headers, pull quotes, and intro copy are easier to build in a word processor than in a transcript editor.
If you’re producing a video, export SRT and check the caption timing on the first minute. Automatic caption breaks occasionally split a sentence awkwardly, and that first minute is where viewers decide whether to keep watching.
What this looks like in practice
A 45-minute one-on-one interview, recorded on two decent mics in a quiet room, processes in minutes. Renaming two speakers takes seconds. A careful cleanup pass with the audio running takes 10 to 20 minutes depending on how much jargon is involved. Export takes one click. Total: under half an hour of your time.
The same interview transcribed by hand takes three to five hours. The difference is entirely in where you spend your attention — reviewing rather than typing.
Listently’s free plan covers 3 transcriptions per day at up to 30 minutes per file, which is enough to test the workflow on a real interview. If your interviews regularly run longer, Pro removes the file-length limit and the daily cap at $8.99/mo billed annually or $13.99/mo billed monthly.
Try it free on your next interview →
Common questions
How long does it take to transcribe an interview? Automatic transcription returns a speaker-labeled transcript in minutes. The manual cleanup pass — fixing names, jargon, and crosstalk errors — typically takes 10 to 20 minutes for a 45-minute interview. Manual transcription from scratch runs three to five hours for the same file.
Can I transcribe an interview recorded on a phone? Yes. Phone recordings in M4A or MP3 work fine as long as the phone was close to the speakers. Place it within arm’s reach, not across a table, and record in a quiet room. Video files from a phone (MP4, MOV) work too — the audio track is extracted automatically.
What if the transcript mislabels who said what? Some crosstalk errors are normal, especially during overlapping speech. Speaker names are editable after transcription, so you can rename labels and correct individual lines. Recording each speaker on a separate mic reduces these errors substantially.
Should I keep filler words in the transcript? Depends on the use. Keep them for legal, academic, or linguistic work where exact speech matters. Remove them for journalism and content — clean verbatim keeps the speaker’s actual words and meaning while cutting “um,” “you know,” and false starts. On Pro, Listently’s “Hide filler words” toggle does this automatically for both the on-screen transcript and exports, without altering the original.