Subtitle Generator — Auto-Generate Subtitles from Video or Audio (Free)
Upload any video or audio file. Get accurate subtitles in 100+ languages with word-level timestamps. Download as SRT, VTT, or TXT — or burn in with custom styling. 30 minutes free, no credit card.
30 min free · SRT / VTT / TXT / burn-in · 100+ languages · Whisper Large-v3
TL;DR
VexaScribe's subtitle generator turns any audio or video into SRT, VTT, TXT, or burned-in subtitles using Whisper Large-v3 in 100+ languages. 30 minutes free, no credit card. Paid plans from $2/mo (200 min) to $20/mo (6,000 min) — roughly $0.60 to $0.20 per hour of media.
Subtitle vs caption — which do you actually want?
Most tools use the terms interchangeably. The technical difference matters for accessibility compliance.
| Format | What it contains | Use when |
|---|---|---|
| Subtitles | Dialogue only — for viewers who can hear the audio | Language translation, sound-off social feeds |
| Closed captions | Dialogue + [music], [applause], speaker labels | ADA / WCAG accessibility compliance |
For the full breakdown of ADA, WCAG 2.1 SC 1.2.2/1.2.4, and DOJ Title II 2024, see What is closed captioning?
How the subtitle generator works
Upload
Drag-and-drop MP3, WAV, M4A, FLAC, MP4, MOV, MKV, or WebM. Files up to 5GB.
Generate
Whisper Large-v3 transcribes and produces word-level timestamps in 100+ languages. Processing takes 20–40% of the media duration.
Download or burn in
Export SRT, VTT, or TXT — or burn subtitles into the video with your font, size, color, and position.
Which format do I need for my platform?
Every platform accepts subtitles differently. Instagram and YouTube Shorts don't support sidecar SRT on native uploads — you have to burn in. Long-form and desktop platforms prefer SRT.
| Platform | Delivery | Notes |
|---|---|---|
| YouTube long-form | SRT | Upload in YouTube Studio → Subtitles. VTT also accepted. |
| YouTube Shorts | Burn-in | Sidecar SRTs are cropped by the vertical UI — burn in for reliability. |
| Instagram Reels | Burn-in | Instagram does not accept SRT upload on Reels (help.instagram.com). Burn in only. |
| TikTok | Burn-in OR SRT base | TikTok caption editor accepts an SRT base; most creators burn in for style control. |
| LinkedIn native video | SRT | Upload SRT in the LinkedIn video composer. |
| Facebook video | SRT | Upload via Creator Studio → Captions. |
| Vimeo | SRT or VTT | Vimeo player supports both; SRT is simplest. |
| Podcast platforms | TXT | Use TXT for show notes and episode transcripts (Apple Podcasts, Spotify). |
| Corporate LMS | SRT or VTT | Most LMS platforms (Cornerstone, Docebo, Canvas) accept both. VTT if you need positioning. |
For a deeper dive on the SRT format, see what is an SRT file? For YouTube specifically, YouTube Studio SRT upload is documented at support.google.com/youtube. Already have an SRT and just need to translate it to another language? Use the dedicated subtitle translator — timestamps preserved, 133 target languages.
Customizing burned-in subtitle style
When you burn subtitles into the video, you control every visual attribute. Common presets for short-form:
- • Font: Inter, Roboto, Montserrat, or Impact (broadcast-style). Custom fonts by upload.
- • Size: 24–48pt on a 1080×1920 vertical frame; 32–60pt for aggressive social styles.
- • Color: White text with a black outline, or brand-color solid fill with drop shadow.
- • Position: Bottom third (default), centered, or upper-third to avoid TikTok/Reels UI overlays.
- • Karaoke / word-highlight: Highlight the active word as it's spoken — matches CapCut and VEED style presets.
- • Background: Transparent, solid, or semi-transparent pill behind each cue.
Language coverage and accuracy
We use OpenAI Whisper Large-v3, which covers 100+ languages. Accuracy varies by language and audio quality. The table below shows word error rate (WER) on the model's benchmark set — lower is better.
| Language | Whisper Large-v3 WER |
|---|---|
| English | ~4.2% |
| Spanish | ~4% |
| French | ~4% |
| Italian | ~5% |
| German | ~5% |
| Portuguese | ~5% |
| Dutch | ~6% |
| Polish | ~6% |
| Russian | ~6% |
| Turkish | ~7% |
| Japanese | ~7% |
| Korean | ~7% |
| Ukrainian | ~7% |
| Mandarin | ~8% |
| Arabic | ~8% |
Source: Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision, arXiv:2212.04356, Table 5. WER measured on FLEURS/CommonVoice benchmark subsets. Real-world audio (background noise, heavy accents, technical vocabulary) will show higher WER.
Editing subtitles after generation
Every subtitle project opens in a full editor with the video preview synced to the cues. You can:
- • Fix misheard words with click-to-edit — timestamps update automatically.
- • Merge or split cues at word boundaries for readability.
- • Rename speakers and preserve speaker labels through SRT export.
- • Adjust cue duration to match speaker delivery (dramatic pauses, fast dialogue).
- • Add non-speech notation like [applause] or [music] for accessibility compliance.
- • Re-export as SRT, VTT, TXT, or a re-burned MP4 without regenerating from scratch.
Bulk generation
Upload up to 50 files at once for batch processing on any paid plan. Useful for course creators subtitling a full module, agencies processing a client backlog, or podcasters generating show-note transcripts across an entire season. Each file lands in your library as an independent project with its own editor and export options.
When to hire a human transcriptionist instead
AI subtitle generators land at 92–96% accuracy on clean audio. For contexts where the remaining 4–8% is unacceptable, use a human service:
- • Broadcast television — FCC accuracy standards require near-perfect captions.
- • Legal proceedings — court transcripts and depositions need certified accuracy.
- • Medical dictation — drug names and procedures where a single misheard word has clinical consequences.
- • Named entities under legal review — patents, contracts, regulatory filings.
A common workflow: run AI generation first (fast + cheap), then hand the SRT to a human editor for a review pass. Services like Rev and Happy Scribe offer human review at $1.50–$2.00 per audio minute — see our comparison of subtitle tools for details.
Frequently asked questions
如何从音频生成字幕?
将您的音频或视频文件通过拖放或文件浏览器上传到VexaScribe。我们的AI转录引擎会处理文件,检测语音内容并生成精确的时间戳,然后创建字幕文件。处理完成后,您可以导出为SRT或VTT格式——两种格式都兼容YouTube、TikTok、LinkedIn以及大多数视频编辑软件。大多数文件只需几分钟即可完成处理。
VexaScribe支持哪些字幕格式?
VexaScribe支持导出SRT(SubRip)和VTT(WebVTT)两种字幕格式。SRT是最广泛支持的格式,兼容YouTube、Premiere Pro、DaVinci Resolve、Final Cut Pro以及大多数社交媒体平台。VTT是HTML5视频播放器使用的Web原生格式,也被YouTube和其他平台所支持。
AI生成的字幕准确度如何?
准确度取决于音频质量、背景噪音和说话者的清晰度。对于背景噪音较少的清晰录音,VexaScribe通常能提供满足专业需求的高准确度。您可以在导出前使用内置编辑器审核和编辑字幕。对于口音较重或包含专业术语的内容,建议进行快速审核。
可以生成不同语言的字幕吗?
可以,VexaScribe支持99种语言的字幕生成,包括英语、西班牙语、法语、德语、葡萄牙语、意大利语、中文、日语、韩语、阿拉伯语、印地语等。系统会自动检测音频中的语言,您也可以手动指定语言以获得最佳效果。
SRT和VTT字幕文件有什么区别?
SRT(SubRip)是使用最广泛的字幕格式——简单、通用,几乎所有视频平台和编辑软件都支持。VTT(WebVTT)是较新的Web原生格式,支持字体颜色和位置等额外样式设置。在大多数情况下,SRT是更稳妥的选择。如果您需要网页播放或自定义样式,请选择VTT。
下载前可以编辑字幕吗?
可以。转录完成后,您可以在VexaScribe的内置编辑器中审核和编辑完整的转录文本。修改词语、调整时间轴、重命名说话者,然后将修正后的版本导出为SRT或VTT格式。这样您无需手动调整时间轴,就能获得专业级别的字幕。
可以上传哪些视频和音频格式?
VexaScribe支持所有常见的音频格式(MP3、WAV、M4A、FLAC、OGG、AAC)和视频格式(MP4、MOV、AVI、MKV、WebM)。对于视频文件,我们会自动提取音频轨道。支持最大5GB的文件。
字幕生成的费用是多少?
字幕生成与转录使用相同的定价。免费试用包含30分钟。付费套餐起步价为每月2美元可使用200分钟(入门版),每月5美元可使用1,000分钟(基础版),每月10美元可使用2,500分钟(专业版),每月20美元可使用6,000分钟(工作室版)。在基础版套餐下,为一个1小时的视频生成字幕大约花费0.30美元。
Comparing tools?
If you're shopping around, we ranked 12 subtitle generators (including ours, honestly at #3) on accuracy, free-tier reality, format support, and per-hour cost.
See the 12-tool ranking →