Subtitle Generator — Auto-Generate Subtitles from Video or Audio (Free)
Upload any video or audio file. Get accurate subtitles in 100+ languages with word-level timestamps. Download as SRT, VTT, or TXT — or burn in with custom styling. 30 minutes free, no credit card.
30 min free · SRT / VTT / TXT / burn-in · 100+ languages · Whisper Large-v3
TL;DR
VexaScribe's subtitle generator turns any audio or video into SRT, VTT, TXT, or burned-in subtitles using Whisper Large-v3 in 100+ languages. 30 minutes free, no credit card. Paid plans from $2/mo (200 min) to $20/mo (6,000 min) — roughly $0.60 to $0.20 per hour of media.
Subtitle vs caption — which do you actually want?
Most tools use the terms interchangeably. The technical difference matters for accessibility compliance.
| Format | What it contains | Use when |
|---|---|---|
| Subtitles | Dialogue only — for viewers who can hear the audio | Language translation, sound-off social feeds |
| Closed captions | Dialogue + [music], [applause], speaker labels | ADA / WCAG accessibility compliance |
For the full breakdown of ADA, WCAG 2.1 SC 1.2.2/1.2.4, and DOJ Title II 2024, see What is closed captioning?
How the subtitle generator works
Upload
Drag-and-drop MP3, WAV, M4A, FLAC, MP4, MOV, MKV, or WebM. Files up to 5GB.
Generate
Whisper Large-v3 transcribes and produces word-level timestamps in 100+ languages. Processing takes 20–40% of the media duration.
Download or burn in
Export SRT, VTT, or TXT — or burn subtitles into the video with your font, size, color, and position.
Which format do I need for my platform?
Every platform accepts subtitles differently. Instagram and YouTube Shorts don't support sidecar SRT on native uploads — you have to burn in. Long-form and desktop platforms prefer SRT.
| Platform | Delivery | Notes |
|---|---|---|
| YouTube long-form | SRT | Upload in YouTube Studio → Subtitles. VTT also accepted. |
| YouTube Shorts | Burn-in | Sidecar SRTs are cropped by the vertical UI — burn in for reliability. |
| Instagram Reels | Burn-in | Instagram does not accept SRT upload on Reels (help.instagram.com). Burn in only. |
| TikTok | Burn-in OR SRT base | TikTok caption editor accepts an SRT base; most creators burn in for style control. |
| LinkedIn native video | SRT | Upload SRT in the LinkedIn video composer. |
| Facebook video | SRT | Upload via Creator Studio → Captions. |
| Vimeo | SRT or VTT | Vimeo player supports both; SRT is simplest. |
| Podcast platforms | TXT | Use TXT for show notes and episode transcripts (Apple Podcasts, Spotify). |
| Corporate LMS | SRT or VTT | Most LMS platforms (Cornerstone, Docebo, Canvas) accept both. VTT if you need positioning. |
For a deeper dive on the SRT format, see what is an SRT file? For YouTube specifically, YouTube Studio SRT upload is documented at support.google.com/youtube. Already have an SRT and just need to translate it to another language? Use the dedicated subtitle translator — timestamps preserved, 133 target languages.
Customizing burned-in subtitle style
When you burn subtitles into the video, you control every visual attribute. Common presets for short-form:
- • Font: Inter, Roboto, Montserrat, or Impact (broadcast-style). Custom fonts by upload.
- • Size: 24–48pt on a 1080×1920 vertical frame; 32–60pt for aggressive social styles.
- • Color: White text with a black outline, or brand-color solid fill with drop shadow.
- • Position: Bottom third (default), centered, or upper-third to avoid TikTok/Reels UI overlays.
- • Karaoke / word-highlight: Highlight the active word as it's spoken — matches CapCut and VEED style presets.
- • Background: Transparent, solid, or semi-transparent pill behind each cue.
Language coverage and accuracy
We use OpenAI Whisper Large-v3, which covers 100+ languages. Accuracy varies by language and audio quality. The table below shows word error rate (WER) on the model's benchmark set — lower is better.
| Language | Whisper Large-v3 WER |
|---|---|
| English | ~4.2% |
| Spanish | ~4% |
| French | ~4% |
| Italian | ~5% |
| German | ~5% |
| Portuguese | ~5% |
| Dutch | ~6% |
| Polish | ~6% |
| Russian | ~6% |
| Turkish | ~7% |
| Japanese | ~7% |
| Korean | ~7% |
| Ukrainian | ~7% |
| Mandarin | ~8% |
| Arabic | ~8% |
Source: Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision, arXiv:2212.04356, Table 5. WER measured on FLEURS/CommonVoice benchmark subsets. Real-world audio (background noise, heavy accents, technical vocabulary) will show higher WER.
Editing subtitles after generation
Every subtitle project opens in a full editor with the video preview synced to the cues. You can:
- • Fix misheard words with click-to-edit — timestamps update automatically.
- • Merge or split cues at word boundaries for readability.
- • Rename speakers and preserve speaker labels through SRT export.
- • Adjust cue duration to match speaker delivery (dramatic pauses, fast dialogue).
- • Add non-speech notation like [applause] or [music] for accessibility compliance.
- • Re-export as SRT, VTT, TXT, or a re-burned MP4 without regenerating from scratch.
Bulk generation
Upload up to 50 files at once for batch processing on any paid plan. Useful for course creators subtitling a full module, agencies processing a client backlog, or podcasters generating show-note transcripts across an entire season. Each file lands in your library as an independent project with its own editor and export options.
When to hire a human transcriptionist instead
AI subtitle generators land at 92–96% accuracy on clean audio. For contexts where the remaining 4–8% is unacceptable, use a human service:
- • Broadcast television — FCC accuracy standards require near-perfect captions.
- • Legal proceedings — court transcripts and depositions need certified accuracy.
- • Medical dictation — drug names and procedures where a single misheard word has clinical consequences.
- • Named entities under legal review — patents, contracts, regulatory filings.
A common workflow: run AI generation first (fast + cheap), then hand the SRT to a human editor for a review pass. Services like Rev and Happy Scribe offer human review at $1.50–$2.00 per audio minute — see our comparison of subtitle tools for details.
Frequently asked questions
Làm thế nào để tạo phụ đề từ âm thanh?
Tải tệp âm thanh hoặc video của bạn lên VexaScribe bằng cách kéo thả hoặc sử dụng trình duyệt tệp. Công cụ phiên âm AI của chúng tôi sẽ xử lý tệp, nhận diện lời nói với dấu thời gian chính xác và tạo ra tệp phụ đề. Sau khi hoàn tất, bạn có thể xuất dưới định dạng SRT hoặc VTT — cả hai đều tương thích với YouTube, TikTok, LinkedIn và hầu hết các trình chỉnh sửa video. Toàn bộ quá trình chỉ mất vài phút cho hầu hết các tệp.
VexaScribe hỗ trợ những định dạng phụ đề nào?
VexaScribe xuất phụ đề ở định dạng SRT (SubRip) và VTT (WebVTT). SRT là định dạng được hỗ trợ rộng rãi nhất và hoạt động với YouTube, Premiere Pro, DaVinci Resolve, Final Cut Pro cùng hầu hết các nền tảng mạng xã hội. VTT là định dạng web gốc được sử dụng bởi trình phát video HTML5 và cũng được chấp nhận bởi YouTube cùng các nền tảng khác.
Phụ đề do AI tạo ra có chính xác không?
Độ chính xác phụ thuộc vào chất lượng âm thanh, tiếng ồn nền và độ rõ ràng của người nói. Đối với các bản ghi âm rõ ràng với ít tiếng ồn nền, VexaScribe thường mang lại độ chính xác cao, phù hợp cho mục đích chuyên nghiệp. Bạn có thể xem lại và chỉnh sửa phụ đề trong trình chỉnh sửa tích hợp trước khi xuất. Đối với nội dung có giọng nặng hoặc thuật ngữ chuyên ngành, nên kiểm tra lại nhanh một lượt.
Tôi có thể tạo phụ đề bằng các ngôn ngữ khác nhau không?
Có, VexaScribe tạo phụ đề bằng 99 ngôn ngữ bao gồm tiếng Anh, tiếng Tây Ban Nha, tiếng Pháp, tiếng Đức, tiếng Bồ Đào Nha, tiếng Ý, tiếng Trung, tiếng Nhật, tiếng Hàn, tiếng Ả Rập, tiếng Hindi và nhiều ngôn ngữ khác. Ngôn ngữ được tự động nhận diện từ âm thanh, hoặc bạn có thể chỉ định thủ công để có kết quả tốt nhất.
Sự khác biệt giữa tệp phụ đề SRT và VTT là gì?
SRT (SubRip) là định dạng phụ đề được sử dụng rộng rãi nhất — đơn giản, phổ biến và được chấp nhận bởi hầu như mọi nền tảng video và trình chỉnh sửa. VTT (WebVTT) là định dạng web mới hơn hỗ trợ thêm các tùy chỉnh kiểu dáng như màu chữ và vị trí hiển thị. Đối với hầu hết các trường hợp sử dụng, SRT là lựa chọn an toàn hơn. Chọn VTT nếu bạn cần phát trên web hoặc tùy chỉnh kiểu dáng.
Tôi có thể chỉnh sửa phụ đề trước khi tải xuống không?
Có. Sau khi phiên âm, bạn có thể xem lại và chỉnh sửa toàn bộ bản ghi trong trình chỉnh sửa tích hợp của VexaScribe. Sửa bất kỳ từ nào, điều chỉnh thời gian, đổi tên người nói, sau đó xuất phiên bản đã chỉnh sửa dưới dạng SRT hoặc VTT. Điều này giúp bạn có phụ đề chất lượng chuyên nghiệp mà không cần phải căn chỉnh thời gian thủ công.
Tôi có thể tải lên những định dạng video và âm thanh nào?
VexaScribe chấp nhận tất cả các định dạng âm thanh phổ biến (MP3, WAV, M4A, FLAC, OGG, AAC) và định dạng video (MP4, MOV, AVI, MKV, WebM). Đối với tệp video, chúng tôi tự động trích xuất rãnh âm thanh. Hỗ trợ tệp có dung lượng lên đến 5GB.
Chi phí tạo phụ đề là bao nhiêu?
Tạo phụ đề sử dụng cùng mức giá với phiên âm. Bản dùng thử miễn phí bao gồm 30 phút. Các gói trả phí bắt đầu từ $2/tháng cho 200 phút (Starter), $5/tháng cho 1.000 phút (Basic), $10/tháng cho 2.500 phút (Pro) và $20/tháng cho 6.000 phút (Studio). Một video dài 1 giờ tốn khoảng $0,30 để tạo phụ đề trên gói Basic.
Comparing tools?
If you're shopping around, we ranked 12 subtitle generators (including ours, honestly at #3) on accuracy, free-tier reality, format support, and per-hour cost.
See the 12-tool ranking →