Home/Interview Transcription
Updated July 2026

Interview Transcription with Speaker Labels — for Journalists, Researchers, and HR

Upload the interview recording — audio or video, up to 5 GB — and get a 92–97% accurate transcript with automatic speaker detection, timestamps, and verbatim/clean-copy options. 30 minutes free. Works for qualitative research, journalism, HR behavioral interviews, and legal/investigative use cases — with the formatting conventions each of those workflows expects.

By VexaScribe Editorial · Updated

Transcribe an Interview in 4 Steps

Same workflow whether you recorded on a phone, headset, Zoom, or DSLR camera.

1

Drop the recording into the uploader

Audio (MP3, WAV, M4A, FLAC, OGG, AAC) or video (MP4, MOV, WEBM, MKV, AVI). Up to 5 GB per file. No pre-conversion needed.

2

Choose verbatim or clean copy

Verbatim retains filler words, false starts, and non-verbal cues — required for qualitative analysis, discourse research, and testimony. Clean copy is edited for readability — better for journalism narrative and internal HR notes.

3

Review speaker labels and proper nouns

Speaker diarization automatically labels interviewer vs interviewee (SPEAKER 1, SPEAKER 2). Rename them once in the editor — every occurrence updates. Spot-check names, organizations, and jargon terms against the audio.

4

Export for your workflow

DOCX for Word/Google Docs, TXT for plain text, PDF for archiving, SRT if the interview is going into video. Line numbers on export for qualitative coding.

Cost reference: a 1-hour interview costs about $0.30 on the Starter plan ($2/mo, 200 minutes) — full comparison against pay-per-file and human interview transcription services on the cost calculator.

Who Uses Interview Transcription

Same tool, different formatting expectations. Match your workflow:

Journalists

Speaker labels for accurate quote attribution, exact-second timestamps for verification, verbatim option for on-the-record quotes vs clean copy for narrative use.

Podcast interview workflow

Qualitative researchers

Coding-friendly speaker labels (SPEAKER 1, SPEAKER 2 — consistent across the corpus), line-numbered export for NVivo/ATLAS.ti/MAXQDA, verbatim retention of pauses and non-verbal cues where discourse analysis requires it.

Academic transcription guide

HR and behavioral interviews

Timestamped record for candidate evaluation and STAR-format coding. Auto-diarization keeps interviewer prompts separate from candidate answers without manual tagging.

Legal, police, and investigative interviews

Word-level timestamps for evidence citation, verbatim mode for testimony-grade accuracy, downloadable transcript with full audit trail. For court-admissible transcription, human transcription may still be required — see the honest AI vs human comparison below.

Legal transcription options

How to Format an Interview Transcript

The formatting decisions that matter for interview transcripts — and what verbatim vs clean-copy conventions actually change:

ElementVerbatimClean copyWhen it matters
Speaker labelsSPEAKER 1:, SPEAKER 2: (default)Interviewer:, [Name]: or first-name attributionRename in the editor once — replaces every occurrence.
TimestampsEvery speaker turn: [00:03:42]Every 5 minutes or at topic changesWord-level available for citation-heavy work.
Filler words (um, uh, like)RetainedRemovedVerbatim required for discourse/conversation analysis.
False starts and repetitionsRetained ("I— I think that—")Cleaned to intended sentenceVerbatim required for linguistic research.
Non-verbal cues[laughter], [pause 3s], [inaudible]Removed unless load-bearingEthnography and clinical research typically retain.
Paragraph breaksPer speaker turnPer speaker turn or topic shiftConsistent breaks make qualitative coding easier.
Line numbersOn exportOn exportRequired by most qualitative coding tools.

Rule of thumb: use verbatim for research, legal, and testimony; clean copy for journalism narrative, HR summaries, and internal knowledge base entries. The transcript is editable in-browser — you can toggle between conventions post-hoc without re-transcribing.

The Three Transcription Levels — True Verbatim, Intelligent Verbatim, Clean

Same recording, three different outputs. Choose the level that matches your job-to-be-done:

LevelExample outputBest for
True verbatim“Um, so like—I think, you know, the thing is— [pause 2s] yeah, we needed to move—faster.”Discourse analysis, conversation analysis, linguistic research, court testimony, legal depositions. Every filler, false start, [pause], [laughter], [overlap] preserved.
Intelligent verbatim“So I think the thing is, yeah, we needed to move faster.”Qualitative research coding (NVivo, MAXQDA, Atlas.ti, Dedoose), UX research, market research, therapist notes. Fillers removed, structure preserved, meaning intact.
Clean copy“The thing is, we needed to move faster.”Journalism, podcasts, published quotes, HR interviews, blog content. Edited for readability; conveys speaker's point, not their exact words.

Which one? Qualitative researchers → intelligent verbatim (most common in published qualitative research per Silverman's Interpreting Qualitative Data and Bryman's Social Research Methods). Journalists and podcasters → clean copy. Legal, court, or discourse analysis → true verbatim. VexaScribe produces intelligent verbatim by default; true verbatim mode retains filler words and non-verbal markers; clean copy is generated on the second pass.

How to Do Verbatim Transcription — the Rules

True verbatim transcription follows conventions from qualitative research methodology and legal transcription. The standard notation:

  • Filler words preserved: um, uh, er, like, you know, I mean — all kept
  • False starts / self-repairs preserved: “I was going—I wanted to say”
  • Non-verbal vocalizations bracketed: [laughter], [sigh], [cough], [throat clear]
  • Pauses marked with duration: [pause 2s], [long pause], [silence]
  • Overlapping speech: [overlapping speech] or Jefferson-style brackets [[ ]]
  • Emphasis marked: WORD in caps for stressed emphasis; underline in DOCX export
  • Inaudible / unclear: [inaudible], [unclear at 12:34], [crosstalk]
  • Non-lexical / paralinguistic: retain hedge markers (mm-hmm, uh-huh) as they carry meaning in interviews
  • Speaker labels consistent: INTERVIEWER, PARTICIPANT (or P1, P2, P3 across a corpus for anonymization)

These conventions are documented in qualitative research methods textbooks (Silverman 2020, Bryman 2016) and Jefferson's conversation-analysis notation system (widely used in discourse analysis). For legal transcription, add certification headers per jurisdiction — see deposition transcription for that workflow.

Export Workflows for NVivo, MAXQDA, Atlas.ti, and Dedoose

The 4 major qualitative analysis tools each have their own import requirements. Here's the honest export recipe for each:

NVivo (QSR International)

  • Best import format: DOCX with paragraph breaks per speaker turn
  • Speaker prefix format: SPEAKER 1: or P1: at start of paragraph
  • Timestamps: optional but useful (every 30 sec or per turn)
  • Export from VexaScribe: DOCX with intelligent verbatim + line numbers

MAXQDA (VERBI Software)

  • Best import format: DOCX or TXT with speaker prefix on each turn
  • MAXQDA's timestamp import supports #00:12:34# inline format
  • For MAXQDA Standard: intelligent verbatim + timestamps every speaker turn
  • Export from VexaScribe: DOCX with MAXQDA timestamp markers

Atlas.ti (Scientific Software Development)

  • Best import format: DOCX with clear speaker headers
  • Atlas.ti auto-detects speakers from SPEAKER N: pattern
  • For synchronized audio coding, keep original audio file paired with transcript
  • Export from VexaScribe: DOCX + keep original MP3/WAV for audio sync

Dedoose (SocioCultural Research Consultants)

  • Cloud-based; imports plain TXT or DOCX
  • Speaker labels via P1: / P2: prefix pattern
  • Line numbering helpful for descriptor tagging
  • Export from VexaScribe: TXT or DOCX with line numbers, intelligent verbatim

All four tools accept the standard intelligent-verbatim DOCX output from VexaScribe. The differences are speaker-prefix format and timestamp syntax. Consistency across your interview corpus matters more than which tool you use.

What the Interview Transcript Looks Like

Clean-copy example from a qualitative interview. Note consistent speaker labels, second-precision timestamps, and paragraph breaks per turn:

[00:00:00] INTERVIEWER: To start, could you tell me a little about your role and how long you've been in it?

[00:00:07] PARTICIPANT: Sure. I've been in product design for about eight years — the last four at my current company. I lead a team of six.

[00:00:16] INTERVIEWER: And when I asked earlier about your day-to-day, you mentioned feeling stretched. Can you walk me through what a typical week looks like?

[00:00:23] PARTICIPANT: Yeah, so — [pause 2s] — I think the hardest part is context-switching. Monday is usually planning, Tuesdays and Wednesdays are heads-down design work, and then Thursdays and Fridays I'm mostly in meetings with engineering and stakeholders.

[00:00:41] INTERVIEWER: When you say "context-switching is the hardest part," what specifically makes it hard?

Verbatim mode retains "yeah, so — [pause 2s] —" and the filler-word pauses shown here. The clean-copy mode above is the default.

Qualitative Interview Transcription for Research

Qualitative research has specific requirements that generic transcription doesn't always meet. What's different:

  • Coding-friendly export: NVivo, ATLAS.ti, and MAXQDA all accept plain-text or DOCX with line numbers. Export line-numbered DOCX for direct import.
  • Consistent speaker attribution: a research corpus of 20 interviews needs the same speaker labeling scheme across all files. SPEAKER 1 / SPEAKER 2 stays consistent; rename to role-based tags (P1, P2, INT) if your codebook expects it.
  • Verbatim requirement for discourse analysis: if you're coding conversation dynamics (turn-taking, overlap, hedging), verbatim mode is not optional. Filler words and pauses are the data.
  • Reflexivity notes stay separate: researcher memos and reflexivity notes shouldn't live in the transcript — keep them in your qualitative software's memo function or a separate log.
  • Anonymization pass: for IRB-sensitive material, run the exported transcript through a redaction step (or use verbatim search-and-replace) before importing to coding software. AI transcription does not automatically redact names.
When AI qualitative transcription is fine: semi-structured interviews with clean headset audio, 2–4 speakers, standard vocabulary. When human transcription is still worth the cost: conversation analysis at turn-construction-unit granularity, dialectal or heavily accented speech in low-resource languages, and any transcript that will appear verbatim in a peer-reviewed journal.

AI Interview Transcription vs Human Transcription

The honest cost/accuracy comparison, without marketing polish:

DimensionAI (this workflow)Human serviceVerdict
Accuracy on clean interview audio (headset mics, quiet room)92–97%99%+Delta rarely justifies the cost gap unless the transcript is court-admissible or medical.
Accuracy on phone / far-field / noisy audio78–90%95–99%Human wins clearly. Consider AI + manual editing pass instead.
Cost per hour$0.20–0.60$60–120100–300× price gap. Human cost stays high even for internal-only interviews.
Turnaround time5–10 minutes12–48 hoursAI beats human transcription for time-to-first-quote by 2 orders of magnitude.
Speaker labelsAutomatic (up to 50 speakers)ManualAI diarization is now reliable for 2–6 speakers. Human still wins on overlapping speech.
Non-verbal cue notationSome ([laughter], [pause])Full ethnographic notationDiscourse and conversation analysis still typically require human transcription.

The middle path most researchers use in 2026: AI transcription (~$0.30/hr) plus a 20–30 minute manual editing pass gets you to ~99% accuracy at 1/100th the wall-clock cost of full human transcription. For step-by-step process guidance, see how to transcribe an interview.

Free Interview Transcription — What 30 Minutes Gets You

Honest version: 30 minutes free covers one standard interview session. It doesn't reset monthly — it's a trial. For ongoing free interview transcription, the actually-free options are:

  • OpenAI Whisper installed locally — free forever, offline, unlimited. Requires Python + a decent GPU. Best for confidential interviews where cloud upload creates compliance risk.
  • Otter.ai free tier — 300 min/month but capped at 30 min per conversation. Fine for short structured interviews, breaks on 60-min narrative ones.
  • TurboScribe free tier — 3 files/day, 30 min cap each.
For researchers running 20+ interviews on a project, the $2/mo Starter plan (200 minutes) covers most single-study corpora at roughly $0.20 per interview. Human interview transcription for the same corpus would cost $1,200–$2,400 at typical $60–$120/hr rates.

Not an Interview? Start on the Right Page

Meeting recording, not an interview — different formatting expectations: Meeting transcription.

Podcast interview episode — show-notes focused: Podcast transcription.

Need a summary, not the full transcript Interview summarizer generates structured key themes and quotes from the same recording.

Step-by-step process guide — workflow-first walkthrough: How to transcribe an interview.

Academic-specific workflow — IRB, dissertation, and thesis transcription: Academic transcription service.

Have the interview recording ready?

Speaker labels, timestamps, verbatim/clean-copy toggle, DOCX/TXT/SRT export. 30 minutes free, no credit card, plans from $2/mo.

Transcribe Interview Free

Frequently Asked Questions

How to format an interview transcript

Two standard conventions: verbatim (retains filler words, false starts, non-verbal cues like [laughter] and [pause 3s]) and clean copy (edited for readability, filler words removed). Both use consistent speaker labels (SPEAKER 1, SPEAKER 2 by default; rename to Interviewer/Interviewee or role tags P1/P2/INT once). Timestamps every speaker turn for verbatim, every 5 minutes for clean copy. Paragraph breaks per speaker turn. Line numbers on export for qualitative coding tools (NVivo, ATLAS.ti, MAXQDA). Verbatim required for research, legal, testimony; clean copy for journalism narrative and internal HR notes.

What's the best interview transcriber for qualitative research?

For qualitative interview transcription, the requirements are: consistent speaker labeling across the corpus (not just within one file), verbatim mode for discourse analysis, line-numbered export for coding software (NVivo/ATLAS.ti/MAXQDA), and 92-97% accuracy on headset audio. VexaScribe supports all four. Otter.ai does verbatim toggle but caps free-tier files at 30 min. Rev's AI mode does verbatim. Human transcription (Rev human, GoTranscript) is still the standard for turn-construction-unit-level conversation analysis and heavily accented speech in low-resource languages — expect $60-120/hr.

How do I transcribe an interview recording?

Four steps: (1) upload the recording (MP3/WAV/M4A audio or MP4/MOV video, up to 5 GB); (2) choose verbatim (research, legal) or clean copy (journalism, HR); (3) review speaker labels — AI diarization auto-labels 2-50 speakers, rename them to INTERVIEWER/PARTICIPANT once; (4) export as DOCX (with line numbers for qualitative coding), TXT (for plain text), PDF (archiving), or SRT (if the interview becomes video). Processing runs 10-30× real-time — a 1-hour interview completes in 2-5 minutes.

Does interview transcription identify who is speaking?

Yes — automatic speaker diarization labels each speaker as SPEAKER 1, SPEAKER 2, etc. by default. Best accuracy with 2-6 distinct voices (interviewer + interviewee is the ideal case). You rename speakers once in the editor and every occurrence updates. Diarization drops accuracy on overlapping speech, very similar-sounding voices, or heavily distant/mono recordings — for those, allocate 15-20 minutes of manual review post-transcription.

How accurate is AI interview transcription?

92-97% on clean interview audio (headset mics, quiet room) per Whisper Large-v3 benchmarks. Accuracy drops on: phone or far-field mic audio (78-90%), heavy accents (85-92%), technical jargon (medical, legal, discipline-specific research vocabulary), and overlapping speakers. For court-admissible transcription or clinical research, add a 20-30 minute manual editing pass to reach 99% at 1/100th the cost of human transcription. Pure human transcription hits 99%+ but costs $60-120/hr with 12-48 hour turnaround.

Is interview transcription free?

Free options: VexaScribe gives 30 minutes on signup (one standard interview). OpenAI Whisper installed locally is free forever, offline, unlimited — the honest choice for confidential interviews. Otter.ai free tier offers 300 min/month but caps at 30 min per conversation. TurboScribe free tier gives 3 files/day, 30-min cap each. For researchers with 20+ interviews, the $2/mo Starter plan covers a full study at ~$0.20 per interview — vs $1,200-$2,400 for the equivalent human transcription corpus.

Can I transcribe qualitative research interviews?

Yes, and it's the fastest-growing use case. Qualitative interview transcription requires: verbatim mode (for discourse analysis), consistent speaker labels across the corpus, line-numbered export for NVivo/ATLAS.ti/MAXQDA, and anonymization pass before coding. VexaScribe handles all four. Human transcription is still preferred for conversation analysis at turn-construction-unit granularity, and for interviews where dialect or accent detail is part of the data.

Can I transcribe police interviews or investigative interviews?

Yes, with important caveats. AI transcription (~92-95% on clean police-interview audio, ~78-88% on typical station-recorded audio) provides a working transcript with word-level timestamps for evidence citation. For court-admissible transcription, human transcription with certification remains the standard in most jurisdictions — see our legal transcription service for that workflow. AI + manual review is common for internal investigations and case-file organization where the transcript won't be introduced as testimony evidence.

What interview formats and file types are supported?

All common audio: MP3, WAV, M4A (iPhone Voice Memos), FLAC, OGG, AAC. All common video: MP4, MOV, WEBM, MKV, AVI. Up to 5 GB per file (roughly 5-6 hours of audio or 720p video). Whether the interview was recorded on Zoom, a phone, a handheld recorder like Zoom H1n, or a DSLR camera, upload directly — no pre-conversion. Rule of thumb: if the file plays in VLC or QuickTime, it transcribes.

Does it work for interviews in other languages?

99 languages supported via Whisper Large-v3 — auto-detected from the audio. Includes Spanish, French, German, Portuguese, Italian, Mandarin, Japanese, Korean, Arabic, Turkish, Hindi, Vietnamese, and dozens more. Cross-language transcription (Spanish audio → English transcript) uses Whisper's translation-to-English mode; for other cross-language pairs, transcribe in source and translate the text separately.

Related Guides