Translate Video to Text — 133 Languages, Timestamps Preserved

Upload video (MP4, MOV, MKV, WebM, AVI). VexaScribe extracts the audio, transcribes in the source language via Whisper Large-v3, then translates the transcript into English (or any of 133 target languages). Export TXT, DOCX, SRT, VTT, or JSON with timestamps and speaker labels preserved through translation. Works with YouTube, TikTok, Instagram, Zoom, and any MP4 file. Free 30 minutes on signup.

133 target languagesTimestamps preservedUp to 5 GB per file
Verified August 12, 2026

How to translate video to text: upload a video file (MP4, MOV, MKV, WebM, AVI — up to 5 GB). VexaScribe extracts the audio track server-side, transcribes in the source language via Whisper Large-v3 (99 supported languages, 5.6% English WER per Radford et al. 2022, arXiv:2212.04356), then translates the transcript into your target from 133 languages using Google Translate (standard tier, free) or our internal translation model (premium tier). Export as TXT, DOCX, SRT (subtitles), VTT (web video), or JSON with timestamps and speaker labels preserved. Free 30 minutes on signup, no credit card.

TL;DR — video-to-text in one flow

  • 1.Upload video. Any format (MP4, MOV, MKV, WebM, AVI). Audio track extracted server-side.
  • 2.Translate. Source language auto-detects. Pick target from 133 languages.
  • 3.Export. TXT, DOCX, SRT (subtitles), VTT (web video), or JSON. Timestamps + speaker labels preserved.

You can search for this as “translate video to text” or “translate video into text” — both preposition forms describe the same job and route to the same workflow.

Text vs dubbed audio. This tool outputs translated text (transcript, SRT, VTT, DOCX) — not a new video file with dubbed audio. For AI-dubbed video (voice-cloned speaker in target language), use ElevenLabs, HeyGen, or Rask AI. That's a different job.

Are you looking to transcribe or translate?

“Translate video to text” can mean two different jobs. Make sure you're asking for the right one:

Your inputYour outputThis isUse this tool?
English videoEnglish textTranscriptionUse video to text instead (faster, no translation stage)
Spanish videoEnglish textTranslation (this page)Yes — you're in the right place
Japanese videoFrench textTranslation (this page)Yes — any source to any of 133 targets
English videoDubbed Spanish audio/videoDubbingUse ElevenLabs, HeyGen, Rask instead

Rule of thumb: if source and target languages differ, this page is right. If they match, use video to text for the transcription-only flow.

How to translate video to text — step-by-step

Step 1 — upload your video file

Drop the video into the upload widget or click to browse. Supported formats: MP4, MOV, WebM, MKV, AVI, WMV, FLV. Maximum 5 GB per file — roughly 2–3 hours of 1080p HD video, or 5+ hours of standard-definition. Audio track is extracted server-side; you don't need to run ffmpeg locally.

Step 2 — auto-detect source language (or set manually)

Source language auto-detects from the first several seconds of audio. Works reliably on 99 supported languages. Override manually for mixed-language content, heavy accents, or very short clips.

Step 3 — pick target language from 133 options

Translation targets cover 133 languages via Google Translate (standard tier) or our internal translation model (premium tier). Common: English for foreign-language sources, or Spanish/French/Chinese/Japanese for English source localization.

Step 4 — export the translated transcript

Download in your chosen format:

  • TXT — plain translated text
  • DOCX — Word document with speaker labels
  • SRT — subtitle file (Premiere, DaVinci, CapCut, YouTube Studio)
  • VTT — WebVTT for HTML5 video
  • JSON — structured data with per-segment timestamps + speaker turns

Video-to-text with timestamps and speaker labels preserved

A common misconception: translation stages break timestamps. They don't — not when the pipeline is designed correctly.

Preservation guarantee

A segment that plays at 00:12:34.500 in the source video appears at 00:12:34.500 in the translated transcript, SRT, or VTT. Only the dialogue text is translated. Sequence numbers, cue timings, and speaker labels (Speaker 1, Speaker 2, up to 50 speakers) survive the translation stage byte-identical to the source.

What this means practically: translated SRT files drop into Premiere Pro, DaVinci Resolve, Final Cut Pro, CapCut, or YouTube Studio without any re-sync work. Multi-speaker videos (podcasts, interviews, panel discussions) preserve speaker attribution across languages — Speaker 1's dialogue is still attributed to Speaker 1 in the translated output.

Platform-specific workflows

Translate YouTube video to text

Two paths: (1) Paste the YouTube URL into YouTube transcript tool with translate-to-target enabled — fastest when captions exist. (2) Download the video, upload here, get translated transcript — works when captions are disabled or accuracy matters. Common workflows: foreign-language YouTube content → English text for research; English YouTube videos → Spanish/French/Chinese/Japanese text for international channel localization.

Translate TikTok video to text

TikTok creators repurposing foreign-language content for English audiences (or vice versa). Download the TikTok clip from the share menu or via TikTok transcript tool, upload here, get translated text back. Short-form video (15-90 seconds) processes in under a minute.

Translate Instagram Reel to text

Same workflow as TikTok. Download the reel or use Instagram transcript tool for URL-paste flow. Great for reel creators localizing content across languages.

Translate Zoom video recording to text

Zoom exports MP4 recordings. Upload the MP4 here, set source to the meeting's primary language, target = English (or your team's working language). Speaker diarization identifies participants. Common for cross-border business meetings, international team standups, cross-cultural user research.

Translate MP4 / MOV / WebM files to text

Any video file you have on disk. Direct upload — VexaScribe extracts the audio server-side. No conversion needed. Formats supported: MP4, MOV, WebM, MKV, AVI, WMV, FLV.

Video speech translator vs video-to-text — same job, different phrasing

If you searched “translate video speech to text”, “translate video voice to text”, or “video speech to text translation” — you're in the right place. The keyword framings are different but describe the same underlying job:

  • “Video to text” = general video-transcript workflow (files, YouTube, Zoom recordings)
  • “Video speech to text” = emphasizes spoken content (as opposed to on-screen text or graphics)
  • “Video voice to text” = same, with voice framing

All three route to this workflow. Whisper Large-v3 handles only spoken content — text overlays, subtitle burn-ins, and on-screen graphics are ignored (they aren't audio-track content).

Language quality tiers for video content

Quality depends on both the transcription stage (Whisper) and the translation stage. Video adds two extra variables: audio track quality (compression, background music, room ambience) and speech clarity (voice-first talking-head videos vs music videos vs on-camera hosting).

Tier 1 — 88–94% word-level accuracy (clean-audio video)

Source language pairs: Spanish, French, German, Portuguese, Italian, Dutch ↔ English. Best on interview footage, talk-format videos, podcasts-with-video, business meetings.

Tier 2 — 82–88% (creator content, first-draft quality)

Source language pairs: Japanese, Korean, Chinese (Mandarin), Arabic, Russian, Hindi, Polish, Turkish ↔ English. Good enough for internal use, content research, and creator workflows. Publication requires bilingual review.

Tier 3 — 75–82% (usable but review-required)

Source language pairs: Vietnamese, Thai, Indonesian, Hebrew, Persian, Ukrainian ↔ English. Auto-translation useful for understanding; manual review needed for external use.

Video-specific quality risks

Music-heavy video (music videos, dance clips): Whisper struggles when speech overlaps music. Expect 20-40% WER drop. Multiple simultaneous speakers: crowd noise, panel-style interruptions. Compressed audio: heavily-compressed video (mobile-shot TikTok/Reels) has lower audio fidelity than studio-recorded.

For per-language WER data across the transcription stage, see how accurate is Whisper. For translation quality tier detail, see translation quality by language pair.

Supported video formats

Video formats

  • .mp4 — most-common container
  • .mov — iPhone recordings, Final Cut
  • .webm — open web video (WebRTC, browser recordings)
  • .mkv — multi-track container
  • .avi / .wmv / .flv — legacy formats supported

File size limits

Maximum 5 GB per file — roughly:

  • 2–3 hours of 1080p HD video
  • 4–5 hours of 720p video
  • 8+ hours of 480p / compressed mobile video

Translate MP4 to text

MP4 is the most-common upload format. Any MP4 works — YouTube downloads, Zoom recordings, iPhone video (converted from MOV), desktop screen recordings, camera footage, GoPro clips. Direct upload, no format conversion needed.

Language pairs — translate video to English or 132 other languages

Most common flow: foreign-language video → English text. Journalists, researchers, and content teams translating Spanish/French/Japanese/Chinese/Arabic video into English drafts.

Reverse (localization) flow: English video → foreign-language text. YouTube creators, brand marketing, e-learning teams localizing English content for international audiences (Spanish/French/German/Portuguese/Chinese/Japanese markets).

Video → Spanish text

LATAM localization

Video → French text

France + French Africa + Quebec

Video → Chinese text

Mainland + Traditional (Taiwan/HK)

Video → Japanese text

Japan market localization

Video → Portuguese text

Brazil + Portugal

Video → German text

DACH region

For audio-only (no video) translation with the same 133-language coverage, use translate audio to text.

Free vs Premium video translation

FeatureFree tierPremium tier
Transcription engineWhisper Large-v3 (Standard)Proprietary higher-accuracy model
Translation engineGoogle TranslateOur internal translation model
Target languages133133
Cost30 min free on signup + included on all plans2 credits per audio min (transcription) + 1 credit per 5,000 chars (translation)
Best forCreator content, YouTube, internal team, first-draftBroadcast subtitles, client-facing content, high-stakes accuracy

When to use another tool

Need dubbed video in target language?

Use ElevenLabs, HeyGen, or Rask AI. Those voice-clone the speaker and generate new video/audio with dubbed dialogue. We output text; they output video.

Just need video editing?

Use Descript, VEED, or Kapwing — they edit video plus offer captions.

Live-streaming transcription?

Use Restream or Otter.ai Live. This tool is batch-only (upload and process); live streams need a real-time system.

Same-language transcription (no translation)?

Use video to text. Faster processing, no translation stage overhead.

Frequently asked questions

How do I translate a video to text?

Upload the video file (MP4, MOV, MKV, WebM, AVI) — up to 5 GB. VexaScribe extracts the audio track server-side, transcribes in the source language (99 supported via Whisper Large-v3), then translates the transcript into your target from 133 languages. Export as TXT, DOCX, SRT (subtitles), VTT (web video), or JSON. Timestamps and speaker labels preserved through translation. Free 30 minutes on signup, no credit card.

How to translate video into text? (preposition variant)

Same workflow — 'translate video to text' and 'translate video into text' describe the same job. Upload the video, pick source and target languages, download the translated transcript. The preposition difference is search-behavior variance, not two different tools.

How to translate a YouTube video to text?

Two paths. (1) Paste the YouTube URL into our YouTube transcript tool with translate-to-English (or any target) enabled — get translated captions directly. (2) Download the YouTube video, upload here, target language of choice, download translated transcript. Path 1 is faster when captions exist; path 2 works even when captions are disabled or accuracy matters more.

Can I translate an MP4 to text?

Yes. MP4 is the most-common video format on VexaScribe. Upload directly — no need to extract audio first, our servers handle that. Also supports MOV, MKV, WebM, AVI, WMV, FLV. Max 5 GB per file (~2-3 hours of 1080p video).

How accurate is video-to-text translation?

Depends on source language and audio quality within the video. Tier 1 sources (Spanish, French, German, Portuguese, Italian ↔ English): 88-94% word-level accuracy on clean audio. Tier 2 (Japanese, Korean, Chinese, Arabic, Russian, Hindi ↔ English): 82-88%. Music-heavy videos (music videos, dance clips) drop accuracy significantly — Whisper struggles when speech and music overlap. Speaker-first videos (interviews, talks, podcasts) work best.

Are timestamps preserved through translation?

Yes. Timestamps stay accurate through the translation stage — a segment that plays at 00:12:34 in the source video will appear at 00:12:34 in the translated transcript, SRT, or VTT export. This means translated SRT files drop into Premiere Pro, DaVinci Resolve, Final Cut, CapCut, or YouTube Studio without re-syncing.

Does it work with TikTok / Instagram videos?

Yes. Download the video (or use the direct URL flow via /tools/tiktok-transcript and /tools/instagram-transcript for URL-paste workflows), upload here, get translated text back. Common use case: creators repurposing foreign-language reels/TikToks for English audiences, or vice versa.

What video formats are supported?

MP4, MOV, WebM, MKV, AVI, WMV, FLV. If your video plays in a modern browser or media player, it will upload. Max 5 GB per file. Audio track is extracted server-side — no local ffmpeg step needed.

Is video-to-text translation free?

Yes on the free tier — 30 minutes of transcription + translation on signup, no credit card. Translation itself is included at no extra cost on every plan. For higher-accuracy translation on client-facing subtitles, premium tier uses our internal translation model at 1 credit per 5,000 characters (typically a fraction of a credit per typical video).

Can I translate videos with multiple speakers?

Yes. Whisper Large-v3 diarization identifies distinct speakers (Speaker 1, Speaker 2, up to 50 speakers per video). Speaker labels persist through the translation stage — you get the translated dialogue attributed to the same speaker across languages. Best on videos with 2-6 clearly-separated speakers; heavy overlap or 20+ speakers drops attribution accuracy.

How long does video translation take?

Roughly equal to video length. A 30-minute video processes in about 30 minutes end-to-end (transcription is the bottleneck; the translation stage is nearly instant on the resulting text). Very short videos (under 5 min) can finish in under a minute. Batch upload up to 50 videos at once via /bulk-transcription.

Can I translate video to Spanish, French, or Chinese text?

Yes. Translation targets cover 133 languages via Google Translate (standard tier) or our internal translation model (premium tier). Any source-language video can translate to any of 133 target languages. Common workflows: English video → Spanish/French/Chinese/Japanese text for international content localization.

Difference between video-to-text transcription and video-to-text translation?

Transcription = same source and target language (English video → English text). Translation = different source and target (Spanish video → English text). This page handles translation. For pure same-language transcription (no translation stage), use /video-to-text — same Whisper Large-v3 engine, faster processing since no translation step. If your source is in the target language, transcription is what you want.

Sources

Page reviewed and accuracy figures verified August 12, 2026.