Home/Audio to SRT
Verified July 27, 2026

Audio to SRT — Convert MP3, WAV, M4A to SRT Subtitle Files (Free)

Upload MP3, WAV, M4A, AAC, OGG, FLAC, WMA, or Opus. VexaScribe transcribes with Whisper Large-v3 and exports a standards-compliant SRT with HH:MM:SS,mmm timecodes and UTF-8 encoding. A 1-hour file processes in 5–10 minutes. 30 minutes free at signup, no card.

By VexaScribe Editorial · Verified

TL;DR

VexaScribe converts MP3, WAV, M4A, AAC, OGG, FLAC, WMA, and Opus audio to SRT subtitle files with Whisper Large-v3 accuracy (~4.2% WER on LibriSpeech clean per arXiv:2212.04356). Three steps: upload, transcribe, export. 1-hour audio processes in 5–10 minutes. Output is standards-compliant SRT with HH:MM:SS,mmm timecodes, UTF-8 encoding, ready for YouTube, Premiere Pro, DaVinci Resolve, or any player.

How to Convert Audio to SRT

End-to-end, a 1-hour audio file becomes a downloadable SRT in about 10–15 minutes of wall-clock time.

1

Upload the audio file

Drag-and-drop MP3, WAV, M4A, AAC, OGG, FLAC, WMA, or Opus into VexaScribe. Files up to 5GB and 4+ hours long are supported. If your audio lives at a public URL (podcast RSS enclosure, direct MP3 link, cloud storage share), paste the URL for the paste-URL flow. iPhone Voice Memos M4A, Zoom cloud recordings, Google Meet recordings — all upload directly with no conversion.

2

Run Whisper Large-v3 transcription

Pick the spoken language (auto-detect handles 100+ languages, but specifying explicitly is more reliable). Enable speaker diarization for multi-speaker audio. VexaScribe runs Whisper Large-v3 (MIT license, released September 2023 by OpenAI) on GPU infrastructure. A 1-hour file typically completes in 5–10 minutes. The dashboard shows progress — you don't need to keep the tab open.

3

Export the SRT

Click Export → SRT. The download is a plain-text .srt file with numbered cues, HH:MM:SS,mmm timecodes, UTF-8 encoding, and speaker prefixes if diarization was enabled. VTT, DOCX, and PDF exports available in the same dialog. Ready for YouTube Studio, Vimeo, LinkedIn Video, or import into Premiere Pro, DaVinci Resolve, or Final Cut Pro.

Supported Audio Formats

Whisper Large-v3 was trained on 680,000 hours of multilingual audio drawn from the open web, so essentially every consumer audio codec works. Upload the format you have — don't pre-convert.

FormatExtensionNotes
MP3.mp3Universal podcast format. 128 kbps or higher: ~0.5–1% WER delta vs WAV. 64 kbps or lower: ~2–4% delta on hard audio.
WAV.wavUncompressed PCM. Baseline reference; highest fidelity but largest file size.
M4A.m4aAAC in an MP4 container. iPhone Voice Memos, Zoom, Google Meet default export. ~0.5–1% WER delta.
AAC.aacRaw AAC audio stream. Common in podcast production and mobile app recording.
OGG.oggVorbis in Ogg container. Open-source, patent-free. ~0.5–1% WER delta.
FLAC.flacFree Lossless Audio Codec. Smaller than WAV, same accuracy.
WMA.wmaWindows Media Audio. Legacy Microsoft codec — accepted but rare in modern workflows.
Opus.opusModern low-latency codec used by Discord, WebRTC, WhatsApp voice notes. Efficient at low bitrates.

Common Use Cases

Podcast episode → captioned social clips

Record the episode, export as MP3 or M4A, upload to VexaScribe, export SRT. The SRT doubles as source for Apple Podcasts / Spotify transcripts, your website transcript page, and burned-in captions for social video clips.

Interview / research audio → timed transcript

Journalists, academic researchers, and podcast interviewers use SRT to pull citable quotes with exact timing. Jump to 00:37:22 in the source to verify a quote — that's the point.

Zoom / Meet / Teams recording → captioned share-out

Zoom exports M4A audio (or MP4 video); Google Meet exports M4A; Teams exports MP4. Upload the audio with speaker diarization on, export the SRT, attach to the recording share link so remote colleagues watch with captions.

Voice memo / dictation → searchable notes

iPhone Voice Memos M4A files upload directly. Export SRT to keep the timestamped structure, or DOCX/TXT for a flat text version. Useful for meeting recap, fieldwork, and note-taking on the go.

Radio show / archived broadcast

Convert broadcast MP3 or WAV archives into SRT for accessibility, search indexing, or podcast repurposing. Diarization separates hosts, guests, and callers.

Oral history / academic research

Historians and qualitative researchers use SRT to anchor quotes to source audio for citation. UTF-8 output handles non-English interviewee names and place names cleanly.

Attaching Audio-Derived SRT to Video

An SRT is a timed subtitle file — useful only when it's paired with playback. Three common patterns:

1. Static-image podcast video. Wrap the podcast audio in a video container with a still image (episode artwork, waveform, or looping brand animation). Any video editor works — Premiere, DaVinci Resolve, CapCut, even FFmpeg. Import the SRT as a caption track or burn it into the video for social platforms that strip caption files (Instagram Reels, some TikTok flows).

2. Separately recorded video (talking-head or interview). If the audio was recorded on a dedicated device (Zoom H6, lav mic into recorder) alongside a video track, the SRT generated from the clean audio track will need slight alignment nudging against the video track — cameras and audio recorders drift by a few frames over long takes. Use your NLE's sync-by-waveform feature to lock the audio track to the video first, then import the SRT.

3. YouTube / Vimeo upload. No editing required — upload the video and SRT separately in the platform's subtitle section. YouTube Studio → Content → Subtitles → Add Language → Upload File. Vimeo Advanced settings during publish. This is the cleanest path for platforms that render SRT natively.

For the reverse case — starting with video and generating SRT directly from the video track — use video to SRT.

Audio Recording Quality — How to Maximise SRT Accuracy

Format choice affects accuracy by under a percentage point. Recording conditions affect it by 15–30 points. The order of impact, from most to least influential:

  • Mic distance — 6–12 inches from the speaker's mouth is the sweet spot. Farther pulls in room noise; closer causes plosive pops that clip.
  • Room acoustics — soft furnishings, closets stuffed with clothes, or a dedicated treated room dramatically outperform reflective spaces (tile bathrooms, empty offices, hard floors with high ceilings).
  • Overlapping speakers — cross-talk is where diarization drops. Enforce turn-taking, use directional mics, or record each speaker on a separate track and merge in post.
  • Sample rate — 44.1 kHz or 48 kHz is standard. Phone-line audio at 8 kHz loses high-frequency detail that speech models rely on; expect 10–15 point WER drop on 8 kHz sources.
  • Bitrate — 128 kbps MP3 is fine; 64 kbps or lower degrades accuracy noticeably, especially on accented speech.

Baseline accuracy figures from the Whisper Large-v3 paper (Radford et al., arXiv:2212.04356) are ~4.2% WER on LibriSpeech clean — a lab benchmark of well-recorded audiobook narration. Real-world figures we measure on user files:

ScenarioAccuracyNotes
Clean single-speaker English podcast (SM7B, PodMic, Blue Yeti)92–95%Best case; broadcast-adjacent quality
2-speaker interview, both mic'd85–92%Speaker attribution usually correct
3+ speaker panel with cross-talk78–88%Diarization drops on overlap
Accented English (non-native or strong regional)75–85%Depends on accent strength
Technical / medical / legal vocabulary70–82%Rare terms substituted with common ones
Noisy environment (café, street, echoic room)65–78%Compression + background noise stack

Full breakdown at how accurate is Whisper.

SRT vs Plain Transcript — Which Do You Actually Want?

Audio produces two kinds of output. Pick the one that matches your use case.

SRT (this page)

Numbered cues with start/end timecodes. Designed to overlay on video during playback. Use for: video captions, podcast video repurposing, jumping to exact quote timing, accessibility compliance for prerecorded video.

Plain transcript (DOCX / TXT / PDF)

Flat text with optional speaker labels but no per-word timing. Use for: reading, editing, LLM input, sharing as a document, blog posts derived from podcasts. See MP3 to text, WAV to text, M4A to text.

Format spec, working code sample, and encoding notes for SRT are covered at what is an SRT file.

When to Hire a Human

Legal depositions, courtroom transcripts, medical records, and regulated accessibility deliverables need human review or human authorship. Whisper Large-v3 gets you 95% of the way there quickly; the last 5% is where mishearings on names, dosages, jurisdictions, and citations create legal risk. For a captioned podcast episode, a lecture recording, a marketing interview, or a YouTube upload, auto-generated SRT with a quick review pass is the right tool.

Audio to SRT FAQ

How do I convert MP3 to SRT?

Upload the MP3 to VexaScribe (30 minutes free at signup), let Whisper Large-v3 transcribe the audio (5–10 minutes for a 1-hour file), then click Export → SRT. The output is a plain-text .srt with numbered cues, HH:MM:SS,mmm timecodes, and UTF-8 encoding — ready for YouTube, Premiere Pro, DaVinci Resolve, or any player that accepts SRT.

How do I convert WAV to SRT?

Same three-step workflow as MP3. WAV is lossless PCM, so it's the highest-fidelity input format — but the accuracy delta versus a good 128 kbps MP3 is under one percentage point on clean audio, so don't bother re-encoding MP3 to WAV before uploading. Upload the file you have. VexaScribe accepts WAV up to 5GB per file.

How do I convert M4A to SRT?

Drag the M4A into VexaScribe. M4A is the default export for iPhone Voice Memos, Zoom cloud recordings, and Google Meet — all of them upload directly with no conversion. Transcription takes 5–10 minutes per hour of audio. Export as SRT when done.

Can I convert audio to SRT for free?

Yes, up to 30 minutes on signup with no card required. Whisper Large-v3 (which powers VexaScribe) is MIT-licensed and free to self-host if you have Python and a GPU, but the DIY route lacks the cue-splitting, editor, and speaker labels. TurboScribe has a limited free tier with ads.

How accurate is audio-to-SRT?

About 92–95% on clean single-speaker English recorded with a proper microphone, per the Whisper Large-v3 paper (Radford et al., arXiv:2212.04356 — ~4.2% WER on LibriSpeech clean). Drops to 75–85% on accented English, 70–82% on technical/medical vocabulary, and 65–78% in noisy environments (café, street, echoic room). Overlapping speakers and phone-line audio (8 kHz sampling) are the main accuracy killers.

How do I attach an SRT file to a video after generating from audio?

Three options depending on your target. (1) Video already exists: import the SRT into Premiere Pro / DaVinci Resolve / Final Cut as a caption track and export the video with subtitles burned in or as a soft-sub track. (2) Static-image video (podcast episode): create a video with a still image using any editor, then attach the SRT. (3) YouTube / Vimeo: upload the video and SRT separately in the platform's subtitle section — no burning required.

What's the difference between audio-to-SRT and audio-to-text?

Audio-to-text produces a flat transcript with no timing — one long text file suitable for reading, editing, or feeding to an LLM. Audio-to-SRT produces a timed subtitle file — numbered cues with start/end timecodes designed to overlay on video during playback. Use SRT when the output needs to sync to video; use plain text when timing isn't required.

Can I convert a voice memo to SRT?

Yes. iPhone Voice Memos exports M4A by default, which uploads directly to VexaScribe. Same three-step workflow. Voice memos are usually single-speaker close-mic recordings, so accuracy lands in the 90%+ range. Useful for turning field notes, dictation, or thinking-out-loud sessions into timed transcripts you can search later.

How long does it take?

About one-fifth the wall time of the audio itself. A 1-hour file typically processes in 5–10 minutes on VexaScribe's GPU infrastructure. Upload takes another minute or two depending on connection; SRT export is instant. You don't need to keep the tab open — the dashboard shows progress and files stay in your workspace.

Convert audio to SRT with VexaScribe (30 min free at signup, no card required) →