Translate Audio to Text — 133 Languages
Upload audio (MP3, WAV, M4A) or video (MP4, MOV). Get an accurate transcript with speaker labels, then translate to any of 133 languages in one flow. Whisper Large-v3 transcription, Google Translate for standard tier, neural MT for premium. Timestamps preserved. Free 30 minutes on signup.
How to translate audio to text: upload an audio or video file (MP3, WAV, M4A, MP4, MOV — up to 5 GB), auto-detect confirms the source language, pick a target from 133 supported languages, click Translate. VexaScribe transcribes with Whisper Large-v3 (5.6% English WER per Radford et al. 2022, arXiv:2212.04356), then translates the transcript via Google Translate. Export as TXT, DOCX, SRT, VTT, or JSON with timestamps and speaker labels preserved. Free 30 minutes on signup, no credit card required.
TL;DR — the 3-step audio translation workflow
- 1.Upload. Any audio (MP3, WAV, M4A, OGG, FLAC, AAC, OPUS) or video (MP4, MOV, MKV, WebM). 5 GB max per file.
- 2.Translate. Source language auto-detects. Pick target from 133 languages. Whisper transcribes, then translation runs on the transcript.
- 3.Export. Download TXT, DOCX, SRT (subtitles), VTT (web video), or JSON. Timestamps and speaker labels preserved.
Text vs dubbed audio. We output translated text (transcript, SRT, VTT, DOCX) — not a new audio file with the speaker's voice speaking the target language. For dubbed audio output (voice-cloned narration in the target language), use ElevenLabs or HeyGen instead. This distinction matters — the top “translate video” SERP is dominated by dubbing tools, which solve a different problem.
How to translate audio to text — step-by-step
The full walkthrough. If you've translated an audio file with any transcription tool before, you can skim this — the flow is standard. If this is your first time, follow the four steps in order.
Step 1 — upload your audio or video file
Drop the file into the upload widget or click to browse. Supported audio: MP3, WAV, M4A, OGG, FLAC, AAC, AMR, OPUS. Supported video: MP4, MOV, WebM, MKV, AVI, WMV, FLV. Maximum 5 GB per file — roughly 5 hours of audio or 2–3 hours of 1080p video.
Video files: the audio track is extracted server-side, so you don't need to run ffmpeg locally. If you only have the audio already extracted, upload that directly — same result.
Step 2 — auto-detect source language (or set manually)
By default, source language is auto-detected from the audio. This works reliably on 99 supported languages. Override manually if the auto-detect is unsure (mixed-language content, heavy accent, or very short clips under 10 seconds).
Common source languages: Spanish, French, German, Portuguese, Italian, Japanese, Korean, Mandarin, Arabic, Russian, Hindi, Turkish, Dutch, Polish, Indonesian, Vietnamese, Thai.
Step 3 — pick target language from 133 options
Translation targets cover 133 languages via Google Translate — every major world language plus many regional and minority languages. Common targets: English, Spanish, French, German, Chinese, Japanese, Portuguese, Arabic, Hindi.
If you want the highest possible accuracy on this pair, switch to premium translation before running (1 credit per 5,000 characters — roughly a fraction of a credit for a 60-minute transcript). Standard Google Translate is fine for creator content and most business use.
Step 4 — export the translated transcript
Once translation completes, download in your chosen format:
- TXT — plain translated text
- DOCX — Word document with speaker labels
- SRT — subtitle file for Premiere, DaVinci, CapCut, YouTube (already have an SRT you just want to translate? use the subtitle translator)
- VTT — WebVTT for HTML5 video players
- JSON — structured data with per-segment timestamps and speaker turns
Worked example
60-minute Spanish podcast episode (MP3, 55 MB) → auto-detect confirms Spanish → target: English → standard translation. Result: ~60-minute transcription pass + ~30-second translation pass = ~60–61 minutes end-to-end. Output: English SRT (7,200 lines average) + English DOCX (~14,000 words) with speaker turns preserved.
Translation quality by language pair — honest tiers
AI audio translation combines two stages: transcription (Whisper Large-v3) plus neural machine translation. Quality depends on both stages, and both stages are stronger on well-resourced language pairs. Marketing pages that claim “99% accuracy” on every language are wrong — here's the honest breakdown.
Tier 1 — 88–94% word-level accuracy on clean audio
Best-supported pairs. Largest Whisper training data + strongest neural MT models.
Pairs: Spanish↔English, French↔English, German↔English, Portuguese↔English, Italian↔English, Dutch↔English.
Tier 2 — 82–88% word-level accuracy
Good enough for creator content, internal use, and first-draft workflows. Expect noticeable misses on idioms and technical vocabulary.
Pairs: Japanese↔English, Korean↔English, Chinese↔English (Mandarin), Arabic↔English, Hindi↔English, Russian↔English, Polish↔English, Turkish↔English.
Tier 3 — 75–82% word-level accuracy
Usable but expect a review pass. Auto-translation is helpful for understanding; manual polishing is required for publication.
Pairs: Vietnamese↔English, Thai↔English, Indonesian↔English, Hebrew↔English, Persian↔English, Ukrainian↔English.
Tier 4 — treat output as first draft only
Low-resource languages (Bengali, Tamil, Urdu, Swahili, Yoruba, Amharic and other African/South Asian languages). Whisper training data is thin, and Google Translate quality drops. Budget for bilingual human review before any external use. For a rough understanding-only pass, still useful.
Whisper Large-v3 English WER benchmark: 5.6% on Common Voice 15 (Radford et al. 2022, arXiv:2212.04356). Non-English WERs vary; see how accurate is Whisper for per-language breakdown.
Translate audio recordings — meetings, interviews, voice memos
Any audio recording works. Common workflows:
Foreign-language interview recording → English transcript
Journalists interviewing sources in Spanish, French, Portuguese, or Arabic get a publishable English draft in under an hour of processing time. Upload the recording, target English, export as DOCX for editing. Speaker labels preserved so quote attribution stays accurate.
Multilingual meeting → English minutes
Cross-border teams recording Zoom/Meet/Teams calls where participants speak different languages. The strongest speaker language auto-detects, transcription runs, then translation to English produces meeting minutes. For mixed-language recordings (heavy code-switching), accuracy drops — consider recording separate language segments.
Voice memo → language-learning practice
Record yourself speaking a target language on iPhone Voice Memos (.m4a) or Android Recorder (.amr, .opus). Upload, translate to your native language, compare with your intent. Useful for pronunciation and grammar self-check.
Lecture recording → foreign-language study notes
Study-abroad students recording lectures in a foreign language get English study notes for review. Upload the lecture recording (typically 60–90 min), target English, export as TXT for pasting into your notes app.
Just want transcription without translation? Use transcribe audio instead — same Whisper Large-v3, no translation step.
Supported audio and video formats
Any common audio or video format works. If your file plays in a modern browser or media player, it will upload.
Audio formats
- .mp3 — universal, podcast standard
- .wav — lossless, high quality
- .m4a — iPhone voice memos, AAC audio
- .ogg / .opus — open format, WhatsApp voice
- .flac — lossless archival
- .aac — broadcast standard
- .amr — Android recorder, mobile calls
Video formats
- .mp4 — most-common video container
- .mov — iPhone recordings, Final Cut
- .webm — open web video
- .mkv — multi-track container
- .avi / .wmv / .flv — legacy formats supported
Translate MP3 audio to English
MP3 is the most-common upload format. Podcast episodes, lecture recordings ripped from YouTube, voice memos exported from phone apps — all work directly. No conversion needed. Source language auto-detects; target = English.
Related: for pure MP3-to-text transcription (no translation), use MP3 to text. Same transcription engine, no translation step.
Translate WAV, M4A, OGG audio
WAV and FLAC are common for lossless recordings (music production sessions, high-quality field recording). M4A is iPhone Voice Memos default. OGG/OPUS is WhatsApp voice message format — upload the .ogg export directly. All produce the same quality translation output as MP3.
Free vs Premium translation — what's included
Two translation tiers. The free tier handles most creator, business, and personal workflows. Premium is for higher-accuracy client-facing content.
| Feature | Free tier | Premium tier |
|---|---|---|
| Translation engine | Google Translate | Our internal translation model |
| Languages supported | 133 | 133 |
| Cost | Included on every plan | 1 credit per 5,000 characters |
| Best for | Creator content, internal use, first draft | Client-facing content, publication, high-stakes accuracy |
| Typical 60-min transcript cost | Free | ~2–3 credits (fraction of monthly allowance) |
Decision framework
Use free tier if: internal team use, social-media captioning, personal understanding, drafts you'll review manually, creator content. Use premium if: paying client deliverable, published article, legal or medical context, subtitles going to broadcast, brand-critical accuracy on names and proper nouns.
Already have a transcript? Translate it in the reader
If you've previously uploaded audio and have the transcript sitting in your VexaScribe library, you don't need to re-run transcription. Open the transcript in the reader view, click Translate in the sidebar, pick your target language. Text updates in-place while speaker labels, timestamps, and formatting stay intact.
3-sentence walkthrough
Open the transcript → click Translate in the reader sidebar → pick target language from 133 options. Translation applies in-place, preserving speaker turns and timestamps. Export the translated version in any supported format (TXT, DOCX, SRT, VTT).
Common transcript translation use cases
- Translate transcript to English from Spanish, French, Japanese, or any other transcription language
- Translate transcript to Spanish for LATAM localization workflows
- Re-translate the transcript after post-processing edits to keep source and translated versions in sync
- Multiple target languages from one source transcript — translate to English, then re-translate the same source to French, German, Japanese without re-uploading audio
Transcripts produced elsewhere (Rev, Otter, Descript, third-party services): you'll need to re-upload the source audio to get an editable in-reader translation. Text-only paste-in translation is not currently supported — the reader requires audio-anchored segments for timestamp preservation.
AI audio translation — how it actually works
Two stages under the hood, each with a Standard and Premium tier. Understanding both explains why quality varies by tier and by language pair.
Stage 1 — transcription (Standard or Premium)
Standard transcription uses OpenAI Whisper Large-v3 — encoder-decoder transformer trained on 680,000 hours of multilingual audio, supporting 99 languages. Uses 1 credit per audio minute. Best for clear audio and standard use (Radford et al. 2022, arXiv:2212.04356).
Premium transcription uses our proprietary higher-accuracy model. Uses 2 credits per audio minute. Superior on quiet, noisy, or accented audio — recommended for professional work where every phrase matters. Same 99-language coverage.
Stage 2 — translation (Standard or Premium)
Standard translation uses Google Translate. Free on every plan, unlimited use, covers all 133 supported languages. Publication-grade on Tier 1 language pairs (Spanish, French, German, Portuguese, Italian ↔ English).
Premium translation uses our internal translation model — higher accuracy on client-facing content, broadcast subtitles, and high-stakes work. Costs 1 credit per 5,000 characters (typically a fraction of a credit per transcript).
Why two stages instead of a single-model translation
Separating transcription and translation gives us: (1) all 133 language targets on the translation side, larger than the 99 transcription languages; (2) source-language transcript available for review before translation; (3) ability to re-translate the same transcript to multiple targets without re-processing audio; (4) independent tier choice — Standard transcription + Premium translation, or vice versa, based on where accuracy matters most for your content.
Honest caveat. Machine translation loses cultural nuance and idiomatic phrasing about 5–10% of the time even on Tier 1 pairs. Technical terminology (medical, legal, engineering jargon) is particularly error-prone. For publication-grade output on high-stakes content, budget for bilingual human review.
Automatic vs manual review — when auto-translation is enough
Auto-translation is fine for…
- Creator content (YouTube, Reels, TikTok subtitles)
- Internal team documents and meeting minutes
- Personal understanding of foreign-language content
- First-draft translations you'll review yourself
- Podcast episode transcripts for repurposing
- Study notes from lectures in a target language
- Interview drafts for later editing
Manual review required for…
- Legal proceedings, depositions, contracts
- Medical records and patient consultations
- Published articles, books, journalism
- Broadcast subtitles going to millions of viewers
- Contracts and financial documents
- Client-facing polished deliverables
- Content where brand names must be exact
Rule of thumb: if a mistranslation would embarrass you or expose you to liability, budget for bilingual review. If it's an internal draft or creator-scale content, auto is fine.
Language pairs — dedicated guides
For deeper coverage of specific language pairs — including regional variant handling (Mexican vs Argentine Spanish, Cantonese vs Mandarin, Brazilian vs European Portuguese), format-specific tips, and worked examples — use the dedicated pages.
Spanish ↔ English
43K/mo cluster. Iberian, Mexican, Argentine, Chilean, Caribbean varieties + WhatsApp voice + voice/speech/audio framing coverage.
French ↔ English
France, Quebecois, Belgian, Swiss, and African French variety coverage. WhatsApp voice + reverse direction (English → French) for Quebec/Africa distribution.
German ↔ English
Hochdeutsch Tier 1 + Austrian + Bavarian/Saxon regional dialects. Honest Swiss German scope (known Whisper weakness). Compound word handling + DACH business.
Chinese ↔ English
Mandarin Tier 1 + honest Cantonese first-draft scope. Simplified vs Traditional character output. WeChat voice messages + tonal audio quality guidance.
Japanese ↔ English
Tokyo standard + Kansai handled well; Tohoku/Kyushu/Okinawan first-draft. Keigo (politeness levels) coverage + anime/manga context + LINE voice notes.
Russian ↔ English
Standard Russian + post-Soviet diaspora varieties (Ukrainian, Belarusian, Baltic, Kazakh Russian) all Tier 1. Telegram voice + independent media use cases.
Video-input translation flow
For video files (MP4, MOV, WebM, MKV) with language translation, use translate video to text — covers intent-split disambiguation (transcribe vs translate), platform workflows (YouTube, TikTok, Instagram, Zoom), and timestamp preservation through translation.
Reverse direction (English → Spanish, English → Japanese, etc.) also works — upload English audio, target any of 132 non-English languages.
When to use another tool
We're not right for every audio-translation scenario. Here's when to reach for something else.
Need dubbed audio in the target language?
Use ElevenLabs, HeyGen, or Rask AI. Those tools voice-clone the original speaker and generate new audio in the target language. We output text (transcript, SRT, VTT, DOCX) — not a new audio file.
Just a short spoken phrase?
Use Google Translate app (voice mode) or DeepL Voice. Free, instant, no file upload needed. Reliable for clips under 5 minutes.
Live speech-to-speech translation?
Use iPhone Live Translate, Google Translate conversation mode, or Samsung Interpreter. Real-time voice-to-voice for travel and in-person conversations. Different workflow from file-based audio translation.
Just text translation (no audio)?
Use DeepL or Google Translate directly. Free for typical use, higher accuracy on text-to-text than audio → text pipelines.
Publication-grade legal or medical translation?
Use a professional translation service with domain-specialist bilingual humans. Machine translation for legal, medical, or high-stakes contexts should be first-draft only. Human translators charge $0.10–0.30 per word for publication-grade output.
Frequently asked questions
How do I translate audio to text?
Upload the audio file (MP3, WAV, M4A, OGG, FLAC, AAC, OPUS) or video file (MP4, MOV, MKV, WebM). VexaScribe auto-detects the source language, transcribes with Whisper Large-v3, then translates the transcript into your chosen target from 133 languages using Google Translate. Export as TXT, DOCX, SRT, VTT, or JSON with timestamps and speaker labels preserved. Free tier: 30 minutes, no credit card.
How do I translate an audio file to English?
Upload the audio file, set the source to auto-detect (or specify Spanish, French, German, Japanese, etc.), and set the target to English. VexaScribe transcribes the source-language audio with Whisper Large-v3 (99 supported languages, 5.6% word error rate on Common Voice 15 English per Radford et al. 2022, arXiv:2212.04356), then translates the transcript to English. Output formats: TXT, DOCX, SRT, VTT. Processing time is roughly equal to audio length — a 30-minute file finishes in about 30 minutes.
How to translate an audio recording?
Same workflow as any audio file: upload the recording (voice memo, interview, meeting, podcast episode, lecture — any source), pick source and target languages, and download the translated transcript. Voice memos from iPhone (.m4a) and Android (.amr, .opus) are supported directly. Meeting recordings from Zoom, Google Meet, or Teams work as MP4 uploads. Longer recordings — interviews, full podcast episodes, 2-hour lectures — get chunked automatically without losing context.
How to translate an MP3 to English?
Upload the MP3, source language auto-detects, target = English. Output is a translated English transcript in TXT, DOCX, SRT, or VTT. Works for MP3s from any source: podcast episodes, lecture recordings, YouTube audio rips, voice memos exported as MP3. Timestamps are preserved so you can jump back to the original audio at any point in the English translation.
How to translate a transcript?
If you already have a transcript on VexaScribe (from a prior upload), open the reader view, click Translate in the sidebar, pick a target language from 133 options — text updates in-place. Speaker labels, timestamps, and formatting stay intact. Export in any supported format. For transcripts produced elsewhere (Rev, Otter, Descript), you'll need to re-upload the source audio to get an editable translated transcript.
Is there a free audio translator?
Yes. VexaScribe includes 30 minutes free translation on signup, no credit card. Translation itself (once transcribed) is unlimited on the free tier — the 30-minute limit applies to transcription. Other free options: Google Translate voice mode (mobile only, short clips), Whisper installed locally (unlimited but requires Python setup + GPU). For creator-scale workflows (5-50 clips/week), a paid tier is usually cheaper per minute than local setup time.
Is audio translation free on all plans?
Yes. Translation is included at no extra cost on every VexaScribe plan including the free trial. Standard translation uses Google Translate — 133 languages, unlimited translations per transcript. Premium neural MT (higher accuracy for client-facing work) costs 1 credit per 5,000 characters — about a fraction of a credit per typical 60-minute transcript.
How to translate Spanish audio to English?
Upload the Spanish audio (MP3, WAV, M4A, MP4). Source auto-detects as Spanish (or set manually), target = English. Spanish → English is the strongest language pair in Whisper's multilingual training — expect ~90-94% word-level accuracy on clean audio. For a full walkthrough with Spanish variant coverage (Mexican, Argentine, Castilian, Caribbean) and video-input workflow, see the dedicated /translate-spanish-audio-to-english page.
Can I transcribe and translate audio in one step?
Yes — that's the core VexaScribe workflow. Upload once, transcription and translation run sequentially in the same page. You don't need two tools or two uploads. The transcript is generated first (so you can review it in the source language if useful), then translated in-place with one click.
How accurate is AI audio translation?
Depends on language pair and audio quality. Tier 1 pairs (Spanish, French, German, Portuguese, Italian ↔ English): 88-94% word-level accuracy on clean audio. Tier 2 (Japanese, Korean, Chinese, Arabic, Hindi, Russian ↔ English): 82-88%. Tier 3 (Vietnamese, Thai, Turkish, Indonesian ↔ English): 75-82%. Tier 4 (low-resource languages): treat output as first draft only. Machine translation loses idioms and cultural nuance about 5-10% of the time — for legal, medical, or publication work, budget for bilingual human review.
What audio and video formats can I translate?
Audio: MP3, WAV, M4A, OGG, FLAC, AAC, AMR, OPUS. Video: MP4, MOV, WebM, MKV, AVI, WMV, FLV. Maximum file size 5 GB — roughly 5 hours of audio or 2-3 hours of 1080p video. Video files: audio track is extracted server-side, no local ffmpeg step needed. YouTube URLs work via the dedicated /tools/youtube-transcript flow.
How many languages are supported for translation?
133 languages for translation via Google Translate — every major world language plus many regional and minority languages. Transcription supports 99 spoken languages via Whisper Large-v3. The translation set is larger than the transcription set because Google Translate covers text-based languages Whisper doesn't recognize acoustically.
Can I translate audio to Spanish, French, or German?
Yes — target any of the 133 supported languages, not just English. Translate English audio to Spanish/French/German/Japanese/etc., or translate any source language to any target. Reverse-direction workflows (English → Spanish, English → Japanese) are common for content localization. VexaScribe runs a two-stage flow for every direction: transcription first (Standard = Whisper Large-v3, Premium = our proprietary model for higher accuracy), then translation (Standard = free Google Translate, Premium = our internal model for professional-grade output).
What's the best AI audio translator?
Depends on the output you need. For translated text (transcript, SRT, VTT, DOCX): VexaScribe, HappyScribe, Notta, VEED, Maestra all offer similar file → translated-text flows. For dubbed audio output (translate + generate voice in target language): ElevenLabs and HeyGen lead — they voice-clone the speaker in the target language. For casual short clips: Google Translate app's voice mode. For unlimited free with technical setup: Whisper installed locally. Match the tool to the output format you actually need.
Is there an audio translation app I need to install?
No install required for VexaScribe — it's browser-based, works on any device with a modern browser (Chrome, Safari, Firefox, Edge). Mobile web works for smaller files. For desktop workflows with 100+ file batches, some users prefer local tools (Whisper CLI, Vovsoft Subtitle Translator) but neither offers our transcription + translation combined flow.
Related tools
Translate Spanish Audio to English
Dedicated language-pair guide with Spanish variant coverage (Mexican, Argentine, Castilian, Caribbean) and video-input workflow.
Transcribe Audio
Just transcription (no translation). Whisper Large-v3, 99 languages, speaker detection.
MP3 to Text
Format-specific transcription for MP3 audio files. Combine with translation for MP3-to-any-language flows.
How Accurate is Whisper
Deep dive on Whisper Large-v3 accuracy per language — the transcription layer under our translation flow.
Subtitle Translator
Translate existing SRT/VTT subtitle files to 133 languages. Timestamps preserved, batch mode up to 50 files.
Multilingual Transcription
99-language transcription with auto-detect. Prerequisite step before translation.
Best Multilingual Transcription Software
10 tools tested across 12 languages. See how VexaScribe compares on non-English accuracy.
Sources
- Radford, A. et al. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356 — source of the Whisper Large-v3 5.6% WER figure on Common Voice 15 English and per-language WER breakdown.
- Google Cloud Translation supported languages — source of the 133-language coverage claim (verified August 2026).
- Mozilla Common Voice 15 dataset — benchmark corpus for Whisper WER evaluation.
- W3C WebVTT specification — standard for the .vtt output format.
Page reviewed and accuracy figures verified August 11, 2026.