Generate animated AI captions for your videos right in your browser — CapCut-style subtitles, no sign-up.
How to use
- Upload your video file to the tool.
- Auto-transcribe the audio or type your own captions, then fine-tune the timing and style.
- Export your captions as SRT or VTT, or burn them directly into the video.
Frequently asked questions
Is it free?
Yes — the Auto Caption Writer is 100% free with no sign-up required.
Do videos upload to a server?
No. Everything runs locally in your browser, so your video never leaves your device.
Which formats can I export?
You can export captions as SRT or VTT subtitle files, or burn the captions directly into your video.
How to Use the Auto Caption Writer
- Upload your video or audio file — or paste a transcript you already have.
- Start transcription. The tool converts speech to text automatically using on-device AI.
- Review and edit the generated captions — fix names, technical terms, or punctuation.
- Style your captions: adjust timing, line breaks, and formatting.
- Export as SRT or VTT subtitle files, or burn captions directly into the video.
Key Features
- Automatic transcription: speech-to-text powered by the Whisper AI model, running locally in your browser.
- SRT & VTT export: standard subtitle formats that work with YouTube, Vimeo, and every major editor.
- Burn-in option: permanently embed styled captions into your video file.
- Editable timing: fine-tune when each caption appears and disappears.
- Private: transcription happens on your device — your audio is never uploaded.
Why Add Captions?
- Accessibility: deaf and hard-of-hearing viewers can follow your content.
- Silent viewing: most social video is watched with sound off — captions keep viewers watching.
- SEO & reach: captioned videos get more engagement and are searchable.
- Clarity: viewers retain more when they read along.
Frequently Asked Questions
How accurate is the transcription?
Very accurate for clear speech — typically 90%+ for well-recorded audio. Always do a quick review pass for names and jargon.
Which languages are supported?
The underlying model supports dozens of languages, with the best accuracy on English and other major languages.
What’s the difference between SRT and VTT?
Both are subtitle files. SRT is the universal standard; VTT adds web styling options. YouTube accepts both.
Is there a file size or length limit?
Processing happens on your device, so limits depend on your device’s memory. Short-form content (under 30 minutes) works smoothly on most devices.
Do my files get uploaded anywhere?
No — transcription and captioning run entirely in your browser. Nothing leaves your device.
How It Works: Under the Hood
Automatic captioning runs in three stages. First, voice activity detection (VAD) scans your audio and marks segments containing speech, skipping silence and background noise. Second, a speech-to-text model (the Whisper architecture family) converts each segment to text: audio is sliced into 30-second windows, transformed into mel spectrograms, and passed through an encoder-decoder transformer that outputs words with start/end timestamps. Third, a segmentation pass splits the raw transcript into caption blocks sized for readability — typically max two lines, ~42 characters per line, each displayed 1–6 seconds. Because the model runs locally in your browser via WebAssembly, your audio is processed on-device: nothing uploads, and timing alignment happens against your file’s actual audio clock, not an estimate.
Real-World Use Cases
- YouTube creators: upload a talking-head video, generate captions, and publish with subtitles — videos watched on mute (up to 85% of mobile views) suddenly retain viewers instead of losing them in 3 seconds.
- Course creators: make every lesson accessible and searchable. Students skim caption text to find the exact minute a concept was explained, cutting support questions.
- Podcasters repurposing to video: transcribe an episode, export SRT, and burn captions onto audiogram clips for TikTok/Reels where sound-off viewing dominates.
- Corporate training teams: compliance often requires captioned training material. Auto-generate the first pass, have a reviewer fix terminology, and ship in a day instead of a week.
- Wedding/event videographers: deliver highlight reels with burned-in subtitles of vows and speeches — clients share these far more than silent montages.
Advanced Tips
- Fix proper nouns first. The model’s weakest point is names, brands, and jargon. Do one review pass focused only on capitalized words before touching anything else — it catches 80% of embarrassing errors.
- Split at natural pauses, not mid-phrase. When editing timing, break lines where the speaker breathes. A caption that reads “…the quarterly revenue / grew forty percent” is dramatically easier to follow than one split mid-thought.
- Use VTT for web, SRT for universal. Export VTT when embedding in an HTML5 player (it supports positioning and styling); use SRT for YouTube/Vimeo uploads where simplicity wins.
- Burn in for social, sidecar for platforms. Instagram and TikTok ignore uploaded SRT files — burn captions into the video for those. YouTube and Vimeo accept sidecar SRT/VTT, which also helps their search indexing.
Common Mistakes to Avoid
- Publishing raw output unchecked. Homophones (“their/there”, “to/too”) and technical terms slip through. One wrong word in a burned-in caption is permanent — always review.
- Wall-of-text captions. A single 100-character line flashing for 2 seconds is unreadable. Keep lines short and on screen long enough to actually read.
- Ignoring speaker changes in interviews. Two-person dialogue needs clear breaks (or labels) per speaker, or viewers can’t follow who said what.
- Wrong format for the destination. Uploading a video with burned-in captions to YouTube wastes the SEO benefit of a separate subtitle track — match the format to the platform.
On-Device AI Means Your Footage Never Leaves Your Computer
Most captioning tools upload your video to a server for processing — a dealbreaker for client work, unreleased content, or anything confidential. This AI caption generator transcribes speech to text with the Whisper model running locally in your browser, so your files stay on your device from upload to export. You get the same automatic transcription workflow without the privacy trade-off, and it works even on a slow connection since nothing needs uploading. When the words matter as much as the visuals, pair it with the hashtag generator to make your captioned clips discoverable.
Polishing Auto Captions: Names, Jargon, and Timing
Automatic transcription gets you 90% of the way there — the last 10% is what separates amateur captions from professional ones. After generating, review the text for proper nouns, brand names, and technical terms the AI may have misheard, then fine-tune the timing so each caption lands with the spoken word. Adjust line breaks so no caption overwhelms the screen, then export: SRT or VTT subtitle files for YouTube and editors, or burned-in captions for feeds where videos autoplay muted. Want help scripting what to say before you record? The AI chat can help you draft it.
Frequently Asked Questions
Can I edit the captions after transcription?
Yes. Review the generated text, fix names, technical terms, or punctuation, adjust timing and line breaks, and then export — nothing is final until you download.
Does it work with audio-only files?
Yes. Upload a video or an audio file, or paste a transcript you already have, and the tool will generate timed captions from it.
Can I style the burned-in captions?
Yes. You can adjust timing, line breaks, and formatting before burning the captions permanently into your video file.