Translate Video to Text — 133 Languages, Timestamps Preserved
Upload video (MP4, MOV, MKV, WebM, AVI). VexaScribe extracts the audio, transcribes in the source language via Whisper Large-v3, then translates the transcript into English (or any of 133 target languages). Export TXT, DOCX, SRT, VTT, or JSON with timestamps and speaker labels preserved through translation. Works with YouTube, TikTok, Instagram, Zoom, and any MP4 file. Free 30 minutes on signup.
How to translate video to text: upload a video file (MP4, MOV, MKV, WebM, AVI — up to 5 GB). VexaScribe extracts the audio track server-side, transcribes in the source language via Whisper Large-v3 (99 supported languages, 5.6% English WER per Radford et al. 2022, arXiv:2212.04356), then translates the transcript into your target from 133 languages using Google Translate (standard tier, free) or our internal translation model (premium tier). Export as TXT, DOCX, SRT (subtitles), VTT (web video), or JSON with timestamps and speaker labels preserved. Free 30 minutes on signup, no credit card.
TL;DR — video-to-text in one flow
- 1.Upload video. Any format (MP4, MOV, MKV, WebM, AVI). Audio track extracted server-side.
- 2.Translate. Source language auto-detects. Pick target from 133 languages.
- 3.Export. TXT, DOCX, SRT (subtitles), VTT (web video), or JSON. Timestamps + speaker labels preserved.
You can search for this as “translate video to text” or “translate video into text” — both preposition forms describe the same job and route to the same workflow.
Text vs dubbed audio. This tool outputs translated text (transcript, SRT, VTT, DOCX) — not a new video file with dubbed audio. For AI-dubbed video (voice-cloned speaker in target language), use ElevenLabs, HeyGen, or Rask AI. That's a different job.
Are you looking to transcribe or translate?
“Translate video to text” can mean two different jobs. Make sure you're asking for the right one:
| Your input | Your output | This is | Use this tool? |
|---|---|---|---|
| English video | English text | Transcription | Use video to text instead (faster, no translation stage) |
| Spanish video | English text | Translation (this page) | Yes — you're in the right place |
| Japanese video | French text | Translation (this page) | Yes — any source to any of 133 targets |
| English video | Dubbed Spanish audio/video | Dubbing | Use ElevenLabs, HeyGen, Rask instead |
Rule of thumb: if source and target languages differ, this page is right. If they match, use video to text for the transcription-only flow.
How to translate video to text — step-by-step
Step 1 — upload your video file
Drop the video into the upload widget or click to browse. Supported formats: MP4, MOV, WebM, MKV, AVI, WMV, FLV. Maximum 5 GB per file — roughly 2–3 hours of 1080p HD video, or 5+ hours of standard-definition. Audio track is extracted server-side; you don't need to run ffmpeg locally.
Step 2 — auto-detect source language (or set manually)
Source language auto-detects from the first several seconds of audio. Works reliably on 99 supported languages. Override manually for mixed-language content, heavy accents, or very short clips.
Step 3 — pick target language from 133 options
Translation targets cover 133 languages via Google Translate (standard tier) or our internal translation model (premium tier). Common: English for foreign-language sources, or Spanish/French/Chinese/Japanese for English source localization.
Step 4 — export the translated transcript
Download in your chosen format:
- TXT — plain translated text
- DOCX — Word document with speaker labels
- SRT — subtitle file (Premiere, DaVinci, CapCut, YouTube Studio)
- VTT — WebVTT for HTML5 video
- JSON — structured data with per-segment timestamps + speaker turns
Video-to-text with timestamps and speaker labels preserved
A common misconception: translation stages break timestamps. They don't — not when the pipeline is designed correctly.
Preservation guarantee
A segment that plays at 00:12:34.500 in the source video appears at 00:12:34.500 in the translated transcript, SRT, or VTT. Only the dialogue text is translated. Sequence numbers, cue timings, and speaker labels (Speaker 1, Speaker 2, up to 50 speakers) survive the translation stage byte-identical to the source.
What this means practically: translated SRT files drop into Premiere Pro, DaVinci Resolve, Final Cut Pro, CapCut, or YouTube Studio without any re-sync work. Multi-speaker videos (podcasts, interviews, panel discussions) preserve speaker attribution across languages — Speaker 1's dialogue is still attributed to Speaker 1 in the translated output.
Platform-specific workflows
Translate YouTube video to text
Two paths: (1) Paste the YouTube URL into YouTube transcript tool with translate-to-target enabled — fastest when captions exist. (2) Download the video, upload here, get translated transcript — works when captions are disabled or accuracy matters. Common workflows: foreign-language YouTube content → English text for research; English YouTube videos → Spanish/French/Chinese/Japanese text for international channel localization.
Translate TikTok video to text
TikTok creators repurposing foreign-language content for English audiences (or vice versa). Download the TikTok clip from the share menu or via TikTok transcript tool, upload here, get translated text back. Short-form video (15-90 seconds) processes in under a minute.
Translate Instagram Reel to text
Same workflow as TikTok. Download the reel or use Instagram transcript tool for URL-paste flow. Great for reel creators localizing content across languages.
Translate Zoom video recording to text
Zoom exports MP4 recordings. Upload the MP4 here, set source to the meeting's primary language, target = English (or your team's working language). Speaker diarization identifies participants. Common for cross-border business meetings, international team standups, cross-cultural user research.
Translate MP4 / MOV / WebM files to text
Any video file you have on disk. Direct upload — VexaScribe extracts the audio server-side. No conversion needed. Formats supported: MP4, MOV, WebM, MKV, AVI, WMV, FLV.
Video speech translator vs video-to-text — same job, different phrasing
If you searched “translate video speech to text”, “translate video voice to text”, or “video speech to text translation” — you're in the right place. The keyword framings are different but describe the same underlying job:
- “Video to text” = general video-transcript workflow (files, YouTube, Zoom recordings)
- “Video speech to text” = emphasizes spoken content (as opposed to on-screen text or graphics)
- “Video voice to text” = same, with voice framing
All three route to this workflow. Whisper Large-v3 handles only spoken content — text overlays, subtitle burn-ins, and on-screen graphics are ignored (they aren't audio-track content).
Language quality tiers for video content
Quality depends on both the transcription stage (Whisper) and the translation stage. Video adds two extra variables: audio track quality (compression, background music, room ambience) and speech clarity (voice-first talking-head videos vs music videos vs on-camera hosting).
Tier 1 — 88–94% word-level accuracy (clean-audio video)
Source language pairs: Spanish, French, German, Portuguese, Italian, Dutch ↔ English. Best on interview footage, talk-format videos, podcasts-with-video, business meetings.
Tier 2 — 82–88% (creator content, first-draft quality)
Source language pairs: Japanese, Korean, Chinese (Mandarin), Arabic, Russian, Hindi, Polish, Turkish ↔ English. Good enough for internal use, content research, and creator workflows. Publication requires bilingual review.
Tier 3 — 75–82% (usable but review-required)
Source language pairs: Vietnamese, Thai, Indonesian, Hebrew, Persian, Ukrainian ↔ English. Auto-translation useful for understanding; manual review needed for external use.
Video-specific quality risks
Music-heavy video (music videos, dance clips): Whisper struggles when speech overlaps music. Expect 20-40% WER drop. Multiple simultaneous speakers: crowd noise, panel-style interruptions. Compressed audio: heavily-compressed video (mobile-shot TikTok/Reels) has lower audio fidelity than studio-recorded.
For per-language WER data across the transcription stage, see how accurate is Whisper. For translation quality tier detail, see translation quality by language pair.
Supported video formats
Video formats
- .mp4 — most-common container
- .mov — iPhone recordings, Final Cut
- .webm — open web video (WebRTC, browser recordings)
- .mkv — multi-track container
- .avi / .wmv / .flv — legacy formats supported
File size limits
Maximum 5 GB per file — roughly:
- 2–3 hours of 1080p HD video
- 4–5 hours of 720p video
- 8+ hours of 480p / compressed mobile video
Translate MP4 to text
MP4 is the most-common upload format. Any MP4 works — YouTube downloads, Zoom recordings, iPhone video (converted from MOV), desktop screen recordings, camera footage, GoPro clips. Direct upload, no format conversion needed.
Language pairs — translate video to English or 132 other languages
Most common flow: foreign-language video → English text. Journalists, researchers, and content teams translating Spanish/French/Japanese/Chinese/Arabic video into English drafts.
Reverse (localization) flow: English video → foreign-language text. YouTube creators, brand marketing, e-learning teams localizing English content for international audiences (Spanish/French/German/Portuguese/Chinese/Japanese markets).
Video → Spanish text
LATAM localization
Video → French text
France + French Africa + Quebec
Video → Chinese text
Mainland + Traditional (Taiwan/HK)
Video → Japanese text
Japan market localization
Video → Portuguese text
Brazil + Portugal
Video → German text
DACH region
For audio-only (no video) translation with the same 133-language coverage, use translate audio to text.
Free vs Premium video translation
| Feature | Free tier | Premium tier |
|---|---|---|
| Transcription engine | Whisper Large-v3 (Standard) | Proprietary higher-accuracy model |
| Translation engine | Google Translate | Our internal translation model |
| Target languages | 133 | 133 |
| Cost | 30 min free on signup + included on all plans | 2 credits per audio min (transcription) + 1 credit per 5,000 chars (translation) |
| Best for | Creator content, YouTube, internal team, first-draft | Broadcast subtitles, client-facing content, high-stakes accuracy |
When to use another tool
Need dubbed video in target language?
Use ElevenLabs, HeyGen, or Rask AI. Those voice-clone the speaker and generate new video/audio with dubbed dialogue. We output text; they output video.
Just need video editing?
Use Descript, VEED, or Kapwing — they edit video plus offer captions.
Live-streaming transcription?
Use Restream or Otter.ai Live. This tool is batch-only (upload and process); live streams need a real-time system.
Same-language transcription (no translation)?
Use video to text. Faster processing, no translation stage overhead.
Frequently asked questions
How do I translate a video to text?
Upload the video file (MP4, MOV, MKV, WebM, AVI) — up to 5 GB. VexaScribe extracts the audio track server-side, transcribes in the source language (99 supported via Whisper Large-v3), then translates the transcript into your target from 133 languages. Export as TXT, DOCX, SRT (subtitles), VTT (web video), or JSON. Timestamps and speaker labels preserved through translation. Free 30 minutes on signup, no credit card.
How to translate video into text? (preposition variant)
Same workflow — 'translate video to text' and 'translate video into text' describe the same job. Upload the video, pick source and target languages, download the translated transcript. The preposition difference is search-behavior variance, not two different tools.
How to translate a YouTube video to text?
Two paths. (1) Paste the YouTube URL into our YouTube transcript tool with translate-to-English (or any target) enabled — get translated captions directly. (2) Download the YouTube video, upload here, target language of choice, download translated transcript. Path 1 is faster when captions exist; path 2 works even when captions are disabled or accuracy matters more.
Can I translate an MP4 to text?
Yes. MP4 is the most-common video format on VexaScribe. Upload directly — no need to extract audio first, our servers handle that. Also supports MOV, MKV, WebM, AVI, WMV, FLV. Max 5 GB per file (~2-3 hours of 1080p video).
How accurate is video-to-text translation?
Depends on source language and audio quality within the video. Tier 1 sources (Spanish, French, German, Portuguese, Italian ↔ English): 88-94% word-level accuracy on clean audio. Tier 2 (Japanese, Korean, Chinese, Arabic, Russian, Hindi ↔ English): 82-88%. Music-heavy videos (music videos, dance clips) drop accuracy significantly — Whisper struggles when speech and music overlap. Speaker-first videos (interviews, talks, podcasts) work best.
Are timestamps preserved through translation?
Yes. Timestamps stay accurate through the translation stage — a segment that plays at 00:12:34 in the source video will appear at 00:12:34 in the translated transcript, SRT, or VTT export. This means translated SRT files drop into Premiere Pro, DaVinci Resolve, Final Cut, CapCut, or YouTube Studio without re-syncing.
Does it work with TikTok / Instagram videos?
Yes. Download the video (or use the direct URL flow via /tools/tiktok-transcript and /tools/instagram-transcript for URL-paste workflows), upload here, get translated text back. Common use case: creators repurposing foreign-language reels/TikToks for English audiences, or vice versa.
What video formats are supported?
MP4, MOV, WebM, MKV, AVI, WMV, FLV. If your video plays in a modern browser or media player, it will upload. Max 5 GB per file. Audio track is extracted server-side — no local ffmpeg step needed.
Is video-to-text translation free?
Yes on the free tier — 30 minutes of transcription + translation on signup, no credit card. Translation itself is included at no extra cost on every plan. For higher-accuracy translation on client-facing subtitles, premium tier uses our internal translation model at 1 credit per 5,000 characters (typically a fraction of a credit per typical video).
Can I translate videos with multiple speakers?
Yes. Whisper Large-v3 diarization identifies distinct speakers (Speaker 1, Speaker 2, up to 50 speakers per video). Speaker labels persist through the translation stage — you get the translated dialogue attributed to the same speaker across languages. Best on videos with 2-6 clearly-separated speakers; heavy overlap or 20+ speakers drops attribution accuracy.
How long does video translation take?
Roughly equal to video length. A 30-minute video processes in about 30 minutes end-to-end (transcription is the bottleneck; the translation stage is nearly instant on the resulting text). Very short videos (under 5 min) can finish in under a minute. Batch upload up to 50 videos at once via /bulk-transcription.
Can I translate video to Spanish, French, or Chinese text?
Yes. Translation targets cover 133 languages via Google Translate (standard tier) or our internal translation model (premium tier). Any source-language video can translate to any of 133 target languages. Common workflows: English video → Spanish/French/Chinese/Japanese text for international content localization.
Difference between video-to-text transcription and video-to-text translation?
Transcription = same source and target language (English video → English text). Translation = different source and target (Spanish video → English text). This page handles translation. For pure same-language transcription (no translation stage), use /video-to-text — same Whisper Large-v3 engine, faster processing since no translation step. If your source is in the target language, transcription is what you want.
Related tools
Video to Text (Same-Language Transcription)
English video → English text. If source language matches target, use this transcription-only flow for faster processing.
Translate Audio to Text
Audio-only sibling — MP3, WAV, M4A input, translated transcript in 133 languages.
Subtitle Translator
Already have an SRT/VTT? Translate the subtitle file directly with timestamps preserved.
YouTube Transcript Tool
Paste YouTube URL directly (no download needed) with translate-to-target enabled.
Bulk Transcription
Translate up to 50 videos per upload for episodic content localization.
Video to SRT
Video file → SRT subtitle file (same-language transcription flow).
Sources
- Radford, A. et al. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356 — source of Whisper Large-v3 per-language WER figures.
- Google Cloud Translation supported languages — 133-language coverage (verified August 2026).
- W3C WebVTT specification — standard for .vtt output format.
Page reviewed and accuracy figures verified August 12, 2026.