Subtitle Generator — Auto-Generate Subtitles from Video or Audio (Free)

Upload any video or audio file. Get accurate subtitles in 100+ languages with word-level timestamps. Download as SRT, VTT, or TXT — or burn in with custom styling. 30 minutes free, no credit card.

30 min free · SRT / VTT / TXT / burn-in · 100+ languages · Whisper Large-v3

TL;DR

VexaScribe's subtitle generator turns any audio or video into SRT, VTT, TXT, or burned-in subtitles using Whisper Large-v3 in 100+ languages. 30 minutes free, no credit card. Paid plans from $2/mo (200 min) to $20/mo (6,000 min) — roughly $0.60 to $0.20 per hour of media.

Subtitle vs caption — which do you actually want?

Most tools use the terms interchangeably. The technical difference matters for accessibility compliance.

FormatWhat it containsUse when
SubtitlesDialogue only — for viewers who can hear the audioLanguage translation, sound-off social feeds
Closed captionsDialogue + [music], [applause], speaker labelsADA / WCAG accessibility compliance

For the full breakdown of ADA, WCAG 2.1 SC 1.2.2/1.2.4, and DOJ Title II 2024, see What is closed captioning?

How the subtitle generator works

Step 1

Upload

Drag-and-drop MP3, WAV, M4A, FLAC, MP4, MOV, MKV, or WebM. Files up to 5GB.

Step 2

Generate

Whisper Large-v3 transcribes and produces word-level timestamps in 100+ languages. Processing takes 20–40% of the media duration.

Step 3

Download or burn in

Export SRT, VTT, or TXT — or burn subtitles into the video with your font, size, color, and position.

Which format do I need for my platform?

Every platform accepts subtitles differently. Instagram and YouTube Shorts don't support sidecar SRT on native uploads — you have to burn in. Long-form and desktop platforms prefer SRT.

PlatformDeliveryNotes
YouTube long-formSRTUpload in YouTube Studio → Subtitles. VTT also accepted.
YouTube ShortsBurn-inSidecar SRTs are cropped by the vertical UI — burn in for reliability.
Instagram ReelsBurn-inInstagram does not accept SRT upload on Reels (help.instagram.com). Burn in only.
TikTokBurn-in OR SRT baseTikTok caption editor accepts an SRT base; most creators burn in for style control.
LinkedIn native videoSRTUpload SRT in the LinkedIn video composer.
Facebook videoSRTUpload via Creator Studio → Captions.
VimeoSRT or VTTVimeo player supports both; SRT is simplest.
Podcast platformsTXTUse TXT for show notes and episode transcripts (Apple Podcasts, Spotify).
Corporate LMSSRT or VTTMost LMS platforms (Cornerstone, Docebo, Canvas) accept both. VTT if you need positioning.

For a deeper dive on the SRT format, see what is an SRT file? For YouTube specifically, YouTube Studio SRT upload is documented at support.google.com/youtube. Already have an SRT and just need to translate it to another language? Use the dedicated subtitle translator — timestamps preserved, 133 target languages.

Customizing burned-in subtitle style

When you burn subtitles into the video, you control every visual attribute. Common presets for short-form:

  • Font: Inter, Roboto, Montserrat, or Impact (broadcast-style). Custom fonts by upload.
  • Size: 24–48pt on a 1080×1920 vertical frame; 32–60pt for aggressive social styles.
  • Color: White text with a black outline, or brand-color solid fill with drop shadow.
  • Position: Bottom third (default), centered, or upper-third to avoid TikTok/Reels UI overlays.
  • Karaoke / word-highlight: Highlight the active word as it's spoken — matches CapCut and VEED style presets.
  • Background: Transparent, solid, or semi-transparent pill behind each cue.

Language coverage and accuracy

We use OpenAI Whisper Large-v3, which covers 100+ languages. Accuracy varies by language and audio quality. The table below shows word error rate (WER) on the model's benchmark set — lower is better.

LanguageWhisper Large-v3 WER
English~4.2%
Spanish~4%
French~4%
Italian~5%
German~5%
Portuguese~5%
Dutch~6%
Polish~6%
Russian~6%
Turkish~7%
Japanese~7%
Korean~7%
Ukrainian~7%
Mandarin~8%
Arabic~8%

Source: Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision, arXiv:2212.04356, Table 5. WER measured on FLEURS/CommonVoice benchmark subsets. Real-world audio (background noise, heavy accents, technical vocabulary) will show higher WER.

Editing subtitles after generation

Every subtitle project opens in a full editor with the video preview synced to the cues. You can:

  • • Fix misheard words with click-to-edit — timestamps update automatically.
  • • Merge or split cues at word boundaries for readability.
  • • Rename speakers and preserve speaker labels through SRT export.
  • • Adjust cue duration to match speaker delivery (dramatic pauses, fast dialogue).
  • • Add non-speech notation like [applause] or [music] for accessibility compliance.
  • • Re-export as SRT, VTT, TXT, or a re-burned MP4 without regenerating from scratch.

Bulk generation

Upload up to 50 files at once for batch processing on any paid plan. Useful for course creators subtitling a full module, agencies processing a client backlog, or podcasters generating show-note transcripts across an entire season. Each file lands in your library as an independent project with its own editor and export options.

When to hire a human transcriptionist instead

AI subtitle generators land at 92–96% accuracy on clean audio. For contexts where the remaining 4–8% is unacceptable, use a human service:

  • Broadcast television — FCC accuracy standards require near-perfect captions.
  • Legal proceedings — court transcripts and depositions need certified accuracy.
  • Medical dictation — drug names and procedures where a single misheard word has clinical consequences.
  • Named entities under legal review — patents, contracts, regulatory filings.

A common workflow: run AI generation first (fast + cheap), then hand the SRT to a human editor for a review pass. Services like Rev and Happy Scribe offer human review at $1.50–$2.00 per audio minute — see our comparison of subtitle tools for details.

Frequently asked questions

如何从音频生成字幕?

将您的音频或视频文件通过拖放或文件浏览器上传到VexaScribe。我们的AI转录引擎会处理文件,检测语音内容并生成精确的时间戳,然后创建字幕文件。处理完成后,您可以导出为SRT或VTT格式——两种格式都兼容YouTube、TikTok、LinkedIn以及大多数视频编辑软件。大多数文件只需几分钟即可完成处理。

VexaScribe支持哪些字幕格式?

VexaScribe支持导出SRT(SubRip)和VTT(WebVTT)两种字幕格式。SRT是最广泛支持的格式,兼容YouTube、Premiere Pro、DaVinci Resolve、Final Cut Pro以及大多数社交媒体平台。VTT是HTML5视频播放器使用的Web原生格式,也被YouTube和其他平台所支持。

AI生成的字幕准确度如何?

准确度取决于音频质量、背景噪音和说话者的清晰度。对于背景噪音较少的清晰录音,VexaScribe通常能提供满足专业需求的高准确度。您可以在导出前使用内置编辑器审核和编辑字幕。对于口音较重或包含专业术语的内容,建议进行快速审核。

可以生成不同语言的字幕吗?

可以,VexaScribe支持99种语言的字幕生成,包括英语、西班牙语、法语、德语、葡萄牙语、意大利语、中文、日语、韩语、阿拉伯语、印地语等。系统会自动检测音频中的语言,您也可以手动指定语言以获得最佳效果。

SRT和VTT字幕文件有什么区别?

SRT(SubRip)是使用最广泛的字幕格式——简单、通用,几乎所有视频平台和编辑软件都支持。VTT(WebVTT)是较新的Web原生格式,支持字体颜色和位置等额外样式设置。在大多数情况下,SRT是更稳妥的选择。如果您需要网页播放或自定义样式,请选择VTT。

下载前可以编辑字幕吗?

可以。转录完成后,您可以在VexaScribe的内置编辑器中审核和编辑完整的转录文本。修改词语、调整时间轴、重命名说话者,然后将修正后的版本导出为SRT或VTT格式。这样您无需手动调整时间轴,就能获得专业级别的字幕。

可以上传哪些视频和音频格式?

VexaScribe支持所有常见的音频格式(MP3、WAV、M4A、FLAC、OGG、AAC)和视频格式(MP4、MOV、AVI、MKV、WebM)。对于视频文件,我们会自动提取音频轨道。支持最大5GB的文件。

字幕生成的费用是多少?

字幕生成与转录使用相同的定价。免费试用包含30分钟。付费套餐起步价为每月2美元可使用200分钟(入门版),每月5美元可使用1,000分钟(基础版),每月10美元可使用2,500分钟(专业版),每月20美元可使用6,000分钟(工作室版)。在基础版套餐下,为一个1小时的视频生成字幕大约花费0.30美元。

Comparing tools?

If you're shopping around, we ranked 12 subtitle generators (including ours, honestly at #3) on accuracy, free-tier reality, format support, and per-hour cost.

See the 12-tool ranking →