Subtitle Generator — Auto-Generate Subtitles from Video or Audio (Free)

Upload any video or audio file. Get accurate subtitles in 100+ languages with word-level timestamps. Download as SRT, VTT, or TXT — or burn in with custom styling. 30 minutes free, no credit card.

30 min free · SRT / VTT / TXT / burn-in · 100+ languages · Whisper Large-v3

TL;DR

VexaScribe's subtitle generator turns any audio or video into SRT, VTT, TXT, or burned-in subtitles using Whisper Large-v3 in 100+ languages. 30 minutes free, no credit card. Paid plans from $2/mo (200 min) to $20/mo (6,000 min) — roughly $0.60 to $0.20 per hour of media.

Subtitle vs caption — which do you actually want?

Most tools use the terms interchangeably. The technical difference matters for accessibility compliance.

FormatWhat it containsUse when
SubtitlesDialogue only — for viewers who can hear the audioLanguage translation, sound-off social feeds
Closed captionsDialogue + [music], [applause], speaker labelsADA / WCAG accessibility compliance

For the full breakdown of ADA, WCAG 2.1 SC 1.2.2/1.2.4, and DOJ Title II 2024, see What is closed captioning?

How the subtitle generator works

Step 1

Upload

Drag-and-drop MP3, WAV, M4A, FLAC, MP4, MOV, MKV, or WebM. Files up to 5GB.

Step 2

Generate

Whisper Large-v3 transcribes and produces word-level timestamps in 100+ languages. Processing takes 20–40% of the media duration.

Step 3

Download or burn in

Export SRT, VTT, or TXT — or burn subtitles into the video with your font, size, color, and position.

Which format do I need for my platform?

Every platform accepts subtitles differently. Instagram and YouTube Shorts don't support sidecar SRT on native uploads — you have to burn in. Long-form and desktop platforms prefer SRT.

PlatformDeliveryNotes
YouTube long-formSRTUpload in YouTube Studio → Subtitles. VTT also accepted.
YouTube ShortsBurn-inSidecar SRTs are cropped by the vertical UI — burn in for reliability.
Instagram ReelsBurn-inInstagram does not accept SRT upload on Reels (help.instagram.com). Burn in only.
TikTokBurn-in OR SRT baseTikTok caption editor accepts an SRT base; most creators burn in for style control.
LinkedIn native videoSRTUpload SRT in the LinkedIn video composer.
Facebook videoSRTUpload via Creator Studio → Captions.
VimeoSRT or VTTVimeo player supports both; SRT is simplest.
Podcast platformsTXTUse TXT for show notes and episode transcripts (Apple Podcasts, Spotify).
Corporate LMSSRT or VTTMost LMS platforms (Cornerstone, Docebo, Canvas) accept both. VTT if you need positioning.

For a deeper dive on the SRT format, see what is an SRT file? For YouTube specifically, YouTube Studio SRT upload is documented at support.google.com/youtube. Already have an SRT and just need to translate it to another language? Use the dedicated subtitle translator — timestamps preserved, 133 target languages.

Customizing burned-in subtitle style

When you burn subtitles into the video, you control every visual attribute. Common presets for short-form:

  • Font: Inter, Roboto, Montserrat, or Impact (broadcast-style). Custom fonts by upload.
  • Size: 24–48pt on a 1080×1920 vertical frame; 32–60pt for aggressive social styles.
  • Color: White text with a black outline, or brand-color solid fill with drop shadow.
  • Position: Bottom third (default), centered, or upper-third to avoid TikTok/Reels UI overlays.
  • Karaoke / word-highlight: Highlight the active word as it's spoken — matches CapCut and VEED style presets.
  • Background: Transparent, solid, or semi-transparent pill behind each cue.

Language coverage and accuracy

We use OpenAI Whisper Large-v3, which covers 100+ languages. Accuracy varies by language and audio quality. The table below shows word error rate (WER) on the model's benchmark set — lower is better.

LanguageWhisper Large-v3 WER
English~4.2%
Spanish~4%
French~4%
Italian~5%
German~5%
Portuguese~5%
Dutch~6%
Polish~6%
Russian~6%
Turkish~7%
Japanese~7%
Korean~7%
Ukrainian~7%
Mandarin~8%
Arabic~8%

Source: Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision, arXiv:2212.04356, Table 5. WER measured on FLEURS/CommonVoice benchmark subsets. Real-world audio (background noise, heavy accents, technical vocabulary) will show higher WER.

Editing subtitles after generation

Every subtitle project opens in a full editor with the video preview synced to the cues. You can:

  • • Fix misheard words with click-to-edit — timestamps update automatically.
  • • Merge or split cues at word boundaries for readability.
  • • Rename speakers and preserve speaker labels through SRT export.
  • • Adjust cue duration to match speaker delivery (dramatic pauses, fast dialogue).
  • • Add non-speech notation like [applause] or [music] for accessibility compliance.
  • • Re-export as SRT, VTT, TXT, or a re-burned MP4 without regenerating from scratch.

Bulk generation

Upload up to 50 files at once for batch processing on any paid plan. Useful for course creators subtitling a full module, agencies processing a client backlog, or podcasters generating show-note transcripts across an entire season. Each file lands in your library as an independent project with its own editor and export options.

When to hire a human transcriptionist instead

AI subtitle generators land at 92–96% accuracy on clean audio. For contexts where the remaining 4–8% is unacceptable, use a human service:

  • Broadcast television — FCC accuracy standards require near-perfect captions.
  • Legal proceedings — court transcripts and depositions need certified accuracy.
  • Medical dictation — drug names and procedures where a single misheard word has clinical consequences.
  • Named entities under legal review — patents, contracts, regulatory filings.

A common workflow: run AI generation first (fast + cheap), then hand the SRT to a human editor for a review pass. Services like Rev and Happy Scribe offer human review at $1.50–$2.00 per audio minute — see our comparison of subtitle tools for details.

Frequently asked questions

چگونه از فایل صوتی زیرنویس تولید کنم؟

فایل صوتی یا ویدیویی خود را با کشیدن و رها کردن یا از طریق مرورگر فایل در VexaScribe بارگذاری کنید. موتور رونویسی هوش مصنوعی ما فایل را پردازش می‌کند، کلمات گفتاری را با برچسب زمانی دقیق شناسایی می‌کند و یک فایل زیرنویس تولید می‌کند. پس از اتمام، خروجی را در قالب SRT یا VTT دریافت کنید — هر دو با YouTube، TikTok، LinkedIn و اکثر ویرایشگرهای ویدیو سازگار هستند. کل فرآیند برای بیشتر فایل‌ها چند دقیقه طول می‌کشد.

VexaScribe از چه فرمت‌های زیرنویسی پشتیبانی می‌کند؟

VexaScribe زیرنویس‌ها را در فرمت‌های SRT (SubRip) و VTT (WebVTT) خروجی می‌دهد. SRT پرکاربردترین فرمت زیرنویس است و با YouTube، Premiere Pro، DaVinci Resolve، Final Cut Pro و اکثر شبکه‌های اجتماعی سازگار است. VTT فرمت بومی وب است که توسط پخش‌کننده‌های ویدیوی HTML5 استفاده می‌شود و همچنین توسط YouTube و سایر پلتفرم‌ها پذیرفته می‌شود.

دقت زیرنویس‌های تولیدشده با هوش مصنوعی چقدر است؟

دقت به کیفیت صدا، نویز پس‌زمینه و وضوح گوینده بستگی دارد. برای ضبط‌های واضح با حداقل نویز پس‌زمینه، VexaScribe معمولاً دقت بالایی ارائه می‌دهد که برای استفاده حرفه‌ای مناسب است. شما می‌توانید قبل از خروجی گرفتن، زیرنویس‌ها را در ویرایشگر داخلی بررسی و ویرایش کنید. برای محتوایی با لهجه‌های سنگین یا اصطلاحات تخصصی، یک بازبینی سریع توصیه می‌شود.

آیا می‌توانم زیرنویس به زبان‌های مختلف تولید کنم؟

بله، VexaScribe زیرنویس را در 99 زبان از جمله انگلیسی، اسپانیایی، فرانسوی، آلمانی، پرتغالی، ایتالیایی، چینی، ژاپنی، کره‌ای، عربی، هندی و بسیاری زبان‌های دیگر تولید می‌کند. زبان به صورت خودکار از صدا شناسایی می‌شود، یا می‌توانید آن را برای بهترین نتیجه به صورت دستی مشخص کنید.

تفاوت بین فایل‌های زیرنویس SRT و VTT چیست؟

SRT (SubRip) پرکاربردترین فرمت زیرنویس است — ساده، جهانی و تقریباً توسط هر پلتفرم ویدیویی و ویرایشگری پذیرفته می‌شود. VTT (WebVTT) فرمت جدیدتر بومی وب است که از قابلیت‌های استایل‌دهی اضافی مانند رنگ فونت و موقعیت‌دهی پشتیبانی می‌کند. برای بیشتر موارد استفاده، SRT انتخاب مطمئن‌تری است. اگر به پخش وب یا استایل‌دهی سفارشی نیاز دارید، VTT را انتخاب کنید.

آیا می‌توانم زیرنویس‌ها را قبل از دانلود ویرایش کنم؟

بله. پس از رونویسی، می‌توانید متن کامل را در ویرایشگر داخلی VexaScribe بررسی و ویرایش کنید. کلمات را اصلاح کنید، زمان‌بندی را تنظیم کنید، نام گویندگان را تغییر دهید و سپس نسخه اصلاح‌شده را به صورت SRT یا VTT خروجی بگیرید. این کار به شما زیرنویس‌هایی با کیفیت حرفه‌ای بدون نیاز به زمان‌بندی دستی می‌دهد.

چه فرمت‌های ویدیویی و صوتی را می‌توانم بارگذاری کنم؟

VexaScribe تمام فرمت‌های رایج صوتی (MP3، WAV، M4A، FLAC، OGG، AAC) و فرمت‌های ویدیویی (MP4، MOV، AVI، MKV، WebM) را می‌پذیرد. برای فایل‌های ویدیویی، ما ترک صوتی را به صورت خودکار استخراج می‌کنیم. فایل‌ها تا حجم 100 مگابایت پشتیبانی می‌شوند.

هزینه تولید زیرنویس چقدر است؟

تولید زیرنویس از همان قیمت‌گذاری رونویسی استفاده می‌کند. دوره آزمایشی رایگان شامل 30 دقیقه است. پلن‌های پولی از 2 دلار/ماه برای 200 دقیقه (Starter)، 5 دلار/ماه برای 1,000 دقیقه (Basic)، 10 دلار/ماه برای 2,500 دقیقه (Pro) و 20 دلار/ماه برای 6,000 دقیقه (Studio) شروع می‌شوند. زیرنویس‌گذاری یک ویدیوی یک ساعته در پلن Basic تقریباً 0.30 دلار هزینه دارد.

Comparing tools?

If you're shopping around, we ranked 12 subtitle generators (including ours, honestly at #3) on accuracy, free-tier reality, format support, and per-hour cost.

See the 12-tool ranking →