Subtitle Generator — Auto-Generate Subtitles from Video or Audio (Free)

Upload any video or audio file. Get accurate subtitles in 100+ languages with word-level timestamps. Download as SRT, VTT, or TXT — or burn in with custom styling. 30 minutes free, no credit card.

30 min free · SRT / VTT / TXT / burn-in · 100+ languages · Whisper Large-v3

TL;DR

VexaScribe's subtitle generator turns any audio or video into SRT, VTT, TXT, or burned-in subtitles using Whisper Large-v3 in 100+ languages. 30 minutes free, no credit card. Paid plans from $2/mo (200 min) to $20/mo (6,000 min) — roughly $0.60 to $0.20 per hour of media.

Subtitle vs caption — which do you actually want?

Most tools use the terms interchangeably. The technical difference matters for accessibility compliance.

FormatWhat it containsUse when
SubtitlesDialogue only — for viewers who can hear the audioLanguage translation, sound-off social feeds
Closed captionsDialogue + [music], [applause], speaker labelsADA / WCAG accessibility compliance

For the full breakdown of ADA, WCAG 2.1 SC 1.2.2/1.2.4, and DOJ Title II 2024, see What is closed captioning?

How the subtitle generator works

Step 1

Upload

Drag-and-drop MP3, WAV, M4A, FLAC, MP4, MOV, MKV, or WebM. Files up to 5GB.

Step 2

Generate

Whisper Large-v3 transcribes and produces word-level timestamps in 100+ languages. Processing takes 20–40% of the media duration.

Step 3

Download or burn in

Export SRT, VTT, or TXT — or burn subtitles into the video with your font, size, color, and position.

Which format do I need for my platform?

Every platform accepts subtitles differently. Instagram and YouTube Shorts don't support sidecar SRT on native uploads — you have to burn in. Long-form and desktop platforms prefer SRT.

PlatformDeliveryNotes
YouTube long-formSRTUpload in YouTube Studio → Subtitles. VTT also accepted.
YouTube ShortsBurn-inSidecar SRTs are cropped by the vertical UI — burn in for reliability.
Instagram ReelsBurn-inInstagram does not accept SRT upload on Reels (help.instagram.com). Burn in only.
TikTokBurn-in OR SRT baseTikTok caption editor accepts an SRT base; most creators burn in for style control.
LinkedIn native videoSRTUpload SRT in the LinkedIn video composer.
Facebook videoSRTUpload via Creator Studio → Captions.
VimeoSRT or VTTVimeo player supports both; SRT is simplest.
Podcast platformsTXTUse TXT for show notes and episode transcripts (Apple Podcasts, Spotify).
Corporate LMSSRT or VTTMost LMS platforms (Cornerstone, Docebo, Canvas) accept both. VTT if you need positioning.

For a deeper dive on the SRT format, see what is an SRT file? For YouTube specifically, YouTube Studio SRT upload is documented at support.google.com/youtube. Already have an SRT and just need to translate it to another language? Use the dedicated subtitle translator — timestamps preserved, 133 target languages.

Customizing burned-in subtitle style

When you burn subtitles into the video, you control every visual attribute. Common presets for short-form:

  • Font: Inter, Roboto, Montserrat, or Impact (broadcast-style). Custom fonts by upload.
  • Size: 24–48pt on a 1080×1920 vertical frame; 32–60pt for aggressive social styles.
  • Color: White text with a black outline, or brand-color solid fill with drop shadow.
  • Position: Bottom third (default), centered, or upper-third to avoid TikTok/Reels UI overlays.
  • Karaoke / word-highlight: Highlight the active word as it's spoken — matches CapCut and VEED style presets.
  • Background: Transparent, solid, or semi-transparent pill behind each cue.

Language coverage and accuracy

We use OpenAI Whisper Large-v3, which covers 100+ languages. Accuracy varies by language and audio quality. The table below shows word error rate (WER) on the model's benchmark set — lower is better.

LanguageWhisper Large-v3 WER
English~4.2%
Spanish~4%
French~4%
Italian~5%
German~5%
Portuguese~5%
Dutch~6%
Polish~6%
Russian~6%
Turkish~7%
Japanese~7%
Korean~7%
Ukrainian~7%
Mandarin~8%
Arabic~8%

Source: Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision, arXiv:2212.04356, Table 5. WER measured on FLEURS/CommonVoice benchmark subsets. Real-world audio (background noise, heavy accents, technical vocabulary) will show higher WER.

Editing subtitles after generation

Every subtitle project opens in a full editor with the video preview synced to the cues. You can:

  • • Fix misheard words with click-to-edit — timestamps update automatically.
  • • Merge or split cues at word boundaries for readability.
  • • Rename speakers and preserve speaker labels through SRT export.
  • • Adjust cue duration to match speaker delivery (dramatic pauses, fast dialogue).
  • • Add non-speech notation like [applause] or [music] for accessibility compliance.
  • • Re-export as SRT, VTT, TXT, or a re-burned MP4 without regenerating from scratch.

Bulk generation

Upload up to 50 files at once for batch processing on any paid plan. Useful for course creators subtitling a full module, agencies processing a client backlog, or podcasters generating show-note transcripts across an entire season. Each file lands in your library as an independent project with its own editor and export options.

When to hire a human transcriptionist instead

AI subtitle generators land at 92–96% accuracy on clean audio. For contexts where the remaining 4–8% is unacceptable, use a human service:

  • Broadcast television — FCC accuracy standards require near-perfect captions.
  • Legal proceedings — court transcripts and depositions need certified accuracy.
  • Medical dictation — drug names and procedures where a single misheard word has clinical consequences.
  • Named entities under legal review — patents, contracts, regulatory filings.

A common workflow: run AI generation first (fast + cheap), then hand the SRT to a human editor for a review pass. Services like Rev and Happy Scribe offer human review at $1.50–$2.00 per audio minute — see our comparison of subtitle tools for details.

Frequently asked questions

오디오에서 자막을 어떻게 생성하나요?

오디오 또는 비디오 파일을 VexaScribe에 드래그 앤 드롭하거나 파일 탐색기를 통해 업로드하세요. AI 전사 엔진이 파일을 처리하고 정확한 타임스탬프와 함께 음성을 인식하여 자막 파일을 생성합니다. 완료되면 SRT 또는 VTT 형식으로 내보낼 수 있으며, 두 형식 모두 YouTube, TikTok, LinkedIn 및 대부분의 영상 편집 프로그램과 호환됩니다. 대부분의 파일은 몇 분 내에 처리가 완료됩니다.

VexaScribe는 어떤 자막 형식을 지원하나요?

VexaScribe는 SRT(SubRip)와 VTT(WebVTT) 형식으로 자막을 내보낼 수 있습니다. SRT는 가장 널리 지원되는 형식으로 YouTube, Premiere Pro, DaVinci Resolve, Final Cut Pro 및 대부분의 소셜 미디어 플랫폼에서 사용할 수 있습니다. VTT는 HTML5 비디오 플레이어에서 사용하는 웹 네이티브 형식이며 YouTube 등 다른 플랫폼에서도 지원됩니다.

AI 생성 자막의 정확도는 어느 정도인가요?

정확도는 오디오 품질, 배경 소음, 화자의 발음 명확성에 따라 달라집니다. 배경 소음이 적은 깨끗한 녹음의 경우 VexaScribe는 전문적인 용도에 적합한 높은 정확도를 제공합니다. 내보내기 전에 내장 편집기에서 자막을 검토하고 수정할 수 있습니다. 강한 억양이나 전문 용어가 포함된 콘텐츠의 경우 간단한 검토를 권장합니다.

다른 언어로 자막을 생성할 수 있나요?

네, VexaScribe는 영어, 스페인어, 프랑스어, 독일어, 포르투갈어, 이탈리아어, 중국어, 일본어, 한국어, 아랍어, 힌디어 등 99개 언어로 자막을 생성할 수 있습니다. 오디오에서 언어가 자동으로 감지되며, 최적의 결과를 위해 수동으로 언어를 지정할 수도 있습니다.

SRT와 VTT 자막 파일의 차이점은 무엇인가요?

SRT(SubRip)는 가장 널리 사용되는 자막 형식으로, 단순하고 범용적이며 거의 모든 비디오 플랫폼과 편집 프로그램에서 지원됩니다. VTT(WebVTT)는 글꼴 색상이나 위치 지정 등 추가 스타일링을 지원하는 최신 웹 네이티브 형식입니다. 대부분의 경우 SRT가 가장 안전한 선택입니다. 웹 재생이나 맞춤 스타일링이 필요한 경우 VTT를 선택하세요.

다운로드 전에 자막을 편집할 수 있나요?

네. 전사가 완료되면 VexaScribe의 내장 편집기에서 전체 텍스트를 검토하고 편집할 수 있습니다. 단어를 수정하고, 타이밍을 조정하고, 화자 이름을 변경한 다음 수정된 버전을 SRT 또는 VTT로 내보낼 수 있습니다. 수동 타이밍 작업 없이 전문적인 품질의 자막을 얻을 수 있습니다.

어떤 비디오 및 오디오 형식을 업로드할 수 있나요?

VexaScribe는 모든 일반적인 오디오 형식(MP3, WAV, M4A, FLAC, OGG, AAC)과 비디오 형식(MP4, MOV, AVI, MKV, WebM)을 지원합니다. 비디오 파일의 경우 오디오 트랙이 자동으로 추출됩니다. 최대 5GB 크기의 파일을 지원합니다.

자막 생성 비용은 얼마인가요?

자막 생성은 전사와 동일한 요금이 적용됩니다. 무료 체험에는 30분이 포함되어 있습니다. 유료 플랜은 월 $2에 200분(Starter), 월 $5에 1,000분(Basic), 월 $10에 2,500분(Pro), 월 $20에 6,000분(Studio)부터 시작합니다. Basic 플랜 기준으로 1시간 영상의 자막 생성 비용은 약 $0.30입니다.

Comparing tools?

If you're shopping around, we ranked 12 subtitle generators (including ours, honestly at #3) on accuracy, free-tier reality, format support, and per-hour cost.

See the 12-tool ranking →