Home/Podcast Transcription
Updated July 2026

Podcast Transcription: Turn Any Episode into a Transcript

Upload the episode — MP3, WAV, M4A, or the video file for a video podcast — and get a 92-97% accurate transcript with multi-host speaker labels, timestamps, and full 3+ hour episode support in 5-25 minutes. Works for solo shows, interview format, panel roundtables, and video podcasts on YouTube. Export as TXT, DOCX, SRT/VTT, or PDF from a single processing pass. 30 minutes free, no credit card, $0.20-$0.60 per hour of audio on paid plans.

By VexaScribe Editorial · Updated

Verified July 28, 2026

In 60 seconds

Podcast transcription converts a spoken episode into text with speaker labels and timestamps. Two paths in 2026: (1) if your podcast is on Spotify or Apple Podcasts, both platforms auto-generate transcripts natively (Spotify since Aug 2023, Apple since Jan 2024) — good for listeners, but no download, no SRT, no speaker labels. (2) For SRT export (YouTube captions), multi-host speaker labels, DOCX for show notes, or podcasts not on Spotify/Apple, upload to a tool like VexaScribe. Real accuracy on clean episodes: 92–95% (Whisper Large V3 + pyannote 4.0, Open ASR Leaderboard). No podcast transcription service reliably hits “99%” despite marketing claims — Rev's human tier is the only 99%+ option, at ~$90 per 1-hour episode. Free tier: 30 min on signup, no card.

Transcribe a Podcast Episode in 4 Steps

Same workflow for solo, interview, and panel formats. Longer episodes take proportionally longer to process — but the steps don't change.

1

Upload the master file

MP3, WAV, M4A, FLAC (audio) or MP4, MOV, MKV (video podcast). Files up to 5 GB — covers roughly 6 hours of high-quality 256 kbps MP3 or 4 hours of 1080p video. If you record on Riverside/StreamYard/Zencastr with per-participant tracks, upload the master or individual tracks — per-track gives cleaner speaker labels.

2

Wait for processing

About 10-15% of audio length. 1-hour episode processes in 5-10 min, 2-hour in 15-20 min, 3-hour in 25-30 min. You can close the browser tab — the transcript saves automatically.

3

Rename speakers and spot-check terms

Auto-labeled speakers (SPEAKER 1, SPEAKER 2) rename to host and guest names once — every occurrence updates. Spot-check proper nouns (company names, technical jargon, guest last names) against the audio timeline. Names are what AI gets wrong most.

4

Export TXT, DOCX, SRT, or VTT

One transcription pass gives you every format. TXT/DOCX for the blog post or show notes; SRT/VTT for YouTube captions on video podcast episodes.

Cost reference: a 1-hour podcast episode costs about $0.30 on the Starter plan ($2/mo, 200 minutes). For a back-catalog of 100 hour-long episodes, the Studio plan ($20/mo, 6,000 minutes) is the honest math — about $0.20/hour all-in. See the cost calculator for the full comparison against pay-per-file and human services.

Podcast Transcript vs Podcast Transcription vs Podcast Transcriber

All the same job, different phrasings. If you landed here from any of these, you're in the right place — the tool below produces the output every phrasing implies.

Search termWhat it actually asks for
Podcast transcriptionThe industry term for the workflow — record the episode, generate the transcript.
Podcast transcriptThe output document. Same thing as 'podcast transcription' — just noun vs process.
Podcast to transcript / podcast to textThe action verb form. Every tool on this page does this.
Podcast transcriberThe AI/service doing the work. "AI podcast transcriber" explicitly rules out human services.
Podcast transcript generatorEmphasizes the AI generator over a human service. Same job, marketing-first phrasing.
Transcribe podcast to textThe explicit request. Same page, same output.

Where's Your Episode? Common Podcast Sources Handled

Spotify, Apple Podcasts, YouTube, or a recording session — the transcription workflow is the same but the download step differs. Six common patterns:

Spotify (own podcast)

Download the master audio from your podcast host (Buzzsprout, Anchor, Libsyn, Transistor, Simplecast) — that's usually higher quality than Spotify's transcoded stream. Upload the MP3 or WAV directly.

Note: Spotify's own auto-generated transcripts (for shows using their creator tools) are typically 80-90% accurate and only exportable inside their platform.

Spotify (someone else's podcast)

You don't have direct MP3 access. Options: (1) use a browser download tool to save the episode locally, then upload; (2) capture audio via a recording tool while the episode plays.

Note: For personal research and journalism this is generally fine; for commercial republishing you need permission from the podcast owner.

Apple Podcasts

Same pattern as Spotify — Apple hosts your feed but the source file lives with your host. Download from your host, upload here.

Note: iOS 17.4+ added on-device transcription for episodes, but export is not straightforward.

YouTube (video podcast)

Paste the YouTube URL directly into our YouTube transcription tool — skips the download step entirely.

Note: For your own uploads, use the master MP4/MOV from your recording tool (Riverside, StreamYard, OBS) for best accuracy.

Riverside / StreamYard / Zencastr

Download the local recording (per-participant tracks) — that's the highest-quality source available. Upload each track separately for cleanest speaker labels, or the mixed track for a single-file workflow.

Note: Per-track recording is the difference between 90% and 98% speaker-label accuracy on multi-guest episodes.

Direct MP3 URL / RSS

Paste the direct episode MP3 URL from the show's RSS feed — many hosts expose these publicly. Skip the download step.

Note: Check the podcast's licensing before using transcripts commercially.

Apple Podcasts & Spotify Native Transcripts — When to Use, When to Skip

Both major platforms added native transcripts for listeners: Spotify in 2023, Apple Podcasts in 2024. If you're a podcast host shipping SEO, show notes, or repurposed content, both have real limitations. Here's the honest picture:

Apple Podcasts native transcripts (2024+)

  • Auto-generated after episode publishes — no host action needed
  • Visible to listeners in Apple Podcasts app with tap-to-jump navigation
  • Language coverage: English + several major languages (verify Apple Podcasts Support for current list)

Limitations for hosts: no editable copy for your website, no SRT/DOCX export, no speaker labels, no control over accuracy, doesn't index on Google. Fine for accessibility in-app; useless for the SEO/show-notes workflow.

Spotify native transcripts (2023+)

  • Auto-generated for select shows using Spotify for Creators
  • Visible on web + mobile Spotify with sync-to-audio
  • Language coverage: English + Spanish + Portuguese (expanding)

Limitations for hosts: Spotify-only visibility, no clean export, editing requires their creator dashboard, no host+guest speaker labeling, no SRT for video podcast, doesn't help your website SEO.

The honest verdict: Native transcripts satisfy in-app listener accessibility. If you also need show notes, blog posts, YouTube captions, or Google-indexable transcript pages, you still need an upload workflow — VexaScribe, Descript, Podsqueeze, Rev, or Whisper local — because native transcripts don't export cleanly. Best pattern: let native transcripts serve the in-app listener; run your own transcription for everything else.

What the Podcast Transcript Actually Looks Like

Clean-copy example from an interview format. Speaker labels, second-precision timestamps, natural paragraph breaks at speaker turns. Verbatim mode retains filler words and [pause] markers.

[00:00:00] HOST: Welcome back to the show. I'm here with Ana Reyes, who runs product at a startup you've probably heard of but I'm gonna let her introduce herself.

[00:00:07] ANA REYES: Thanks for having me. So — I lead product at Kestrel. We're the API most people don't know they're using — payment reconciliation across ~40 markets.

[00:00:18] HOST: Right, and you joined how long ago?

[00:00:20] ANA REYES: Uh, coming up on four years. I was employee 12, we're at about 240 now.

[00:00:26] HOST: What surprised you the most about scaling from that size?

[00:00:30] ANA REYES: [pause] Honestly? How much of the job becomes internal comms. When there were 12 of us I could literally walk to anyone's desk. At 240 you need process, and the failure mode isn't that people don't know what to do — it's that they build the wrong thing because you didn't communicate the why.

The generated transcript is editable in-browser — fix a proper noun once, it stays fixed across the whole file. SRT export takes the same speaker turns and slices them at subtitle-appropriate boundaries for video podcast uploads to YouTube/Vimeo.

Podcast Transcript vs Show Notes — What's the Difference?

Same source (the episode), different deliverables:

DimensionTranscriptShow notesVerdict
Length8,000-15,000 words for a 1-hour interview300-800 wordsDifferent outputs, different jobs
FormatVerbatim, with timestamps and speaker labelsCurated summary + timestamped chapter markers + links mentionedShow notes are written FROM the transcript
Job to be doneSEO, accessibility, content repurposing, research, quote sourcingGive listeners a scan-able overview + links to referenced materialBoth matter for a serious podcast
AutomationAI does it well (92-97% on clean audio)AI can draft, but human editorial pass neededTranscript = fully automatable; show notes = mostly automatable
Publish locationFull page on your podcast website (huge SEO value)Episode description on Spotify/Apple + episode page on your sitePublish both; they capture different audiences

Workflow: transcribe the episode with this tool → get the full transcript → optionally feed the transcript into our podcast summarizer to draft show notes with chapter markers and quotes → publish both on your episode page (transcript for SEO, show notes for scan-ability).

Speaker Labels for Multi-Host Podcasts — The Honest Reality

Podcasts with 2 co-hosts plus rotating guests are the hardest case for automatic speaker diarization. Nobody in this SERP gives you real numbers. Here are ours, based on internal testing of Whisper Large V3 + pyannote 4.0 across ~200 podcast episodes.

FormatSpeaker-label accuracyManual cleanup needed?
Solo host (monologue)~99%None
Solo host + 1 guest (interview)~95–98%2–3 mislabels per hour
2 co-hosts (no guest)~95%Rename "Speaker 1"/"Speaker 2" once; propagates
2 hosts + 1 guest~92%5–8 mislabels per hour, mostly near overlap
Panel: 3–4 speakers~80–90%Significant — expect 15–30 min cleanup per hour
Chaotic roundtable: 5+ speakers overlapping~75%Consider recording separate tracks per host

The hardest cases: same-gender voices with similar tone, thick simultaneous cross-talk, and heavily processed audio (compressors making all voices sound similar). Best practice: if your show records each host on a separate track (Riverside, Squadcast, StreamYard, Zencastr all do this), upload the mixed master — the underlying audio still helps diarization. If you record everything through one shared mic, expect the middle rows of this table. Renaming labels in the editor propagates to the full transcript in one click.

What to Do with the Transcript — Show Notes, YouTube, RSS, Blog

Getting the transcript is step one. Getting value from it is the rest of the work. Four practical outputs:

📝 Show notes

Paste the transcript into a template with these blocks: 2–3 sentence hook, 5–10 timestamped chapter markers, 3 pull-quotes with speaker attribution, list of links/books/tools mentioned.

Automated version: feed the transcript into our podcast summarizer for the whole draft in one step.

📹 YouTube captions (SRT)

Video podcast on YouTube? Export as SRT (Subtitle → SRT in the editor). Upload alongside the video in YouTube Studio → Subtitles → Add. YouTube also has its own auto-captions but Whisper accuracy on cleanly recorded podcast audio typically beats them, especially for names.

See also: audio to SRT workflow.

📡 RSS transcript block

The Podcast Index 2.0 namespace supports <podcast:transcript> — a tag on your RSS feed pointing to the transcript file. Podcast apps that support it (Podverse, Overcast Boost) show the transcript in-app. Your host either supports adding this (Buzzsprout, Transistor, RSS.com do) or you edit the RSS manually. Confidence: medium — adoption varies by app.

🔍 Blog post / SEO

Publish the full transcript on your episode page. Google indexes every word. Pacific Content and Edison Research have documented full-transcript shows seeing meaningful organic traffic lift over 6–12 months (typically 2–5×) versus audio-only competitors. Add a table of contents linking to key moments (using the timestamps you already have).

What Podcasters Actually Do with Transcripts

The transcript is the source material. Six most common uses:

SEO — publish the transcript on your episode page
Podcast audio is invisible to Google; the transcript makes every spoken word indexable. Podcasters who publish full transcripts report meaningful organic search lift over 6-12 months (Pacific Content, Edison Research, and independent case studies from Descript / Riverside all document this).
Show notes and episode summaries
Feed the transcript into a summary tool to draft chapter markers, key quotes, and a 2-3 paragraph episode description. Faster than writing show notes from memory.
Repurposing — blog posts, LinkedIn, Twitter/X threads
Pull specific quotes with attribution and timestamps directly from the transcript. Every 1-hour episode is 3-5 blog posts of raw material.
Video podcast subtitles (SRT/VTT)
Same transcription pass gives you the transcript + the SRT file for YouTube captions. No re-upload needed.
Accessibility
The CDC estimates ~15% of US adults have some degree of hearing loss. Publishing transcripts makes your show accessible to that audience and generally to non-native listeners.
Research and journalism
Transcribing someone else's podcast for research is generally fine under fair-use principles. Upload the MP3 or paste the YouTube URL.

Free Podcast Transcription — What Actually Works

Honest version: for one short episode, a free tool is genuinely fine. For a weekly show or a back catalog, free tiers run out fast. The realistic options for podcasters:

  • VexaScribe 30-min trial — one 30-min episode or two 15-min clips. One-time, doesn't reset. Full feature access on trial (speaker labels, SRT export, no watermark).
  • TurboScribe free — 3 files/day, 30 min each. Fine for daily short-form; useless for long-form interviews.
  • Otter.ai free — 300 min/month, capped at 30 min per file. Optimized for meetings; long-form episodes get truncated.
  • OpenAI Whisper installed locally — free forever, unlimited, offline. Requires Python + GPU for reasonable speed. Best for privacy-sensitive shows (therapy, legal, healthcare podcasts).
  • Descript free — 60 min/month with watermark. Useful if you're also going to edit the podcast in Descript; otherwise the export is limited.
  • Spotify auto-transcripts — if your show is hosted on Spotify Podcasters, they generate them, but export is not straightforward.
Break-even math for podcasters: weekly hour-long show = ~4 hours/month of transcription. Free tiers cover it barely; VexaScribe Starter ($2/mo, 200 minutes) covers it 12× over with room for a back catalog. For a full back-catalog project (100+ episodes), Studio at $20/mo for one month is the honest math — about $0.20/hour all-in.

Not Just a Podcast? Start on the Right Page

Video podcast on YouTube — paste the URL directly: YouTube transcription. Faster than downloading the MP4 first.

Interview transcription (not podcast format) Interview transcription has verbatim/clean-copy options and coding-friendly export for qualitative research.

You want show notes, not the full transcript Podcast summarizer takes the transcript and generates structured chapter markers, key quotes, and a 2-3 paragraph summary.

100+ back-catalog episodes to process Bulk transcription handles 50-file parallel uploads on paid plans.

Comparing tools before committing Best podcast transcription tools ranks 10+ options against each other.

Have an episode ready?

Drop the file — multi-host speaker labels, timestamps, 99 languages, TXT/DOCX/SRT/VTT export. 30 minutes free, no credit card, plans from $2/mo.

Transcribe a Podcast Episode Free

Frequently Asked Questions

What's the best podcast transcription tool?

Depends on your workflow. For most independent podcasters and small networks who want a clean transcript with multi-host speaker labels, VexaScribe gives 30 minutes free on signup, then $2–$20/month for higher volume — at the $20 Studio tier, that works out to roughly $0.20 per hour of audio. Otter has a generous free tier (300 min/month) but is meeting-recording-first; Descript is excellent if you also want a podcast editor in the same tool (different product category — they own that space); Riverside bundles recording + transcription at $24+/month if you also need to record remotely. Rev's human transcription is the most accurate (~99%) but costs ~$90 per 1-hour episode — only worth it for high-stakes work. For pure cost-per-minute at scale, install OpenAI Whisper locally and pay $0.

Should I publish podcast transcripts for SEO?

Yes — and most podcasters don't realize the size of the opportunity. Podcast audio is invisible to Google search by default; the only thing search engines can index is your episode title and description. Publishing the transcript turns every spoken word into searchable text. Pacific Content and Edison Research have repeatedly documented that shows publishing full transcripts see meaningful organic traffic lift compared to audio-only shows — typical reports range from 2–5× organic search growth over 6–12 months. Bonus: accessibility (the CDC estimates ~15% of US adults have some degree of hearing loss) and international audience (transcripts can be translated, audio can't).

How accurate are multi-speaker podcast transcripts?

Whisper-based speaker diarization (which VexaScribe uses) is most accurate with 2–4 distinct voices. Realistic accuracy by speaker count: 2 speakers (typical solo + guest) → 95%+ label accuracy; 3–4 speakers (host + 2–3 guests) → 90–95%; 5–6 speakers (panel format) → 80–90%; 7+ speakers (chaotic roundtables) → requires manual cleanup. The hardest cases: same-gender voices with similar tone, and any segment with overlapping speech. Best practice for podcasters: after the first transcription pass, rename "Speaker 1" → host name, "Speaker 2" → guest name, then save the named pattern for future episodes.

Can it handle 2- or 3-hour episodes?

Yes — long-form is increasingly common (Joe Rogan, Tim Ferriss, Lex Fridman, Acquired all run 2–4+ hour episodes). VexaScribe processes long episodes as a single file with no need to split. Realistic timing: 1-hour episode ≈ 5–10 min to process; 2-hour ≈ 15–20 min; 3-hour ≈ 25–30 min. File size cap is 5 GB per upload, which covers roughly 6 hours of high-quality 256 kbps MP3 or about 4 hours of 1080p video podcast. Most free transcription tools cap at 25 MB (~30 minutes of audio) — a real constraint for the long-form format.

Does it work for video podcasts (YouTube format)?

Yes. Upload the MP4/MOV directly — VexaScribe extracts audio internally. No need to convert. If your video podcast lives on YouTube and you don't have the source file, our YouTube transcription tool accepts video URLs directly. For Riverside recordings (high-quality WAV + MP4), use either file. The transcript output is the same; the SRT export is useful if you're also uploading the video to YouTube/Vimeo and want captions.

How long does it take to transcribe a podcast episode?

About 10–15% of audio length on AI tools: 1-hour episode → ~5–10 min, 2-hour → ~15–20 min, 3-hour → ~25–30 min. Processing happens server-side, so you can close the browser tab and come back when it's done — the transcript saves automatically. Human transcription services (Rev, GoTranscript) take 4–24 hours regardless of episode length.

What's the difference between a transcript and show notes?

They're different deliverables for different jobs. A transcript is the full literal text of everything said in the episode — typically 8,000–15,000 words for a 1-hour interview, mostly used for SEO (publishing on your website), accessibility, content repurposing, and research. Show notes are a curated summary: 2–4 paragraphs of highlights, a timestamped list of topics or chapter markers, and links to anything mentioned — typically 300–800 words, written FROM the transcript after the episode is done. VexaScribe produces the transcript; show notes are something you write (or generate with our summary tool) from the transcript as raw material.

Can I transcribe someone else's podcast for research?

Yes — for personal research, journalism, academic study, competitive analysis, or quote sourcing, transcribing a podcast you don't own is generally fine under fair-use principles (specifics vary by jurisdiction). You can either upload an MP3/MP4 of the episode you've saved locally, or use our YouTube transcription tool if it's a video podcast on YouTube. For commercial republishing (e.g., publishing transcripts of someone else's podcast on your own site as content), you'd need permission from the podcast creator — the transcript itself can be a derivative work for copyright purposes.

What audio and video formats work for podcast transcription?

Audio formats from any podcast host or recording app: MP3 (Buzzsprout, Anchor, Libsyn exports), WAV/AIFF (studio sessions in Hindenburg, Pro Tools, Audacity, Reaper), M4A (iPhone/QuickTime field recordings), FLAC, OGG, AAC. Video formats for video podcasts: MP4, MOV, MKV, WEBM. Mix audio + video in the same workflow without conversion — VexaScribe handles audio extraction automatically.

What's the cheapest way to transcribe a back-catalog of 100+ episodes?

For 100 typical 1-hour podcast episodes (~6,000 minutes total), the math: VexaScribe Studio at $20/month covers it — that's roughly $0.20/hour or $0.003/minute, all-in flat pricing. Deepgram API is roughly $0.22/hour pay-as-you-go at base rates ($22 for the full batch) but requires developer setup. Rev human transcription at $1.25–$1.99/min would cost $7,500–$11,940 for the same 100 episodes — only worth it if you need legal-grade accuracy. Whisper installed locally is $0 if you have a GPU machine and the patience for batch scripting. For most podcasters with a back-catalog, VexaScribe Studio for one month is the simplest path. See our bulk transcription page for the parallel-upload workflow.

Related Guides