SRT/VTT Subtitles from AI Voiceover: Zero-ASR Timing
Updated September 3, 2026
TL;DR: when the audio is generated from a script, the script is the ground truth — so LeapFun's Studio builds subtitles from the narration timeline, not from speech recognition. Paragraph boundaries are exact; sentence timing inside a paragraph is proportional. No misheard words, no external service, no cost. Export SRT for most editors and platforms, VTT for the web.
Why not just transcribe the audio?
Auto-captions listen to the output and guess. They mishear names, drop short words, and add a processing step. A voiceover generated from text doesn't need guessing: every word and its order are known. LeapFun records where each paragraph starts and how long it runs when it merges the chapter, splits the paragraph text into sentences, and assigns times inside the paragraph by character share.
How accurate is the timing?
- Paragraph boundaries: exact, from the merge timeline.
- Sentence boundaries inside a paragraph: proportional to character count — typically within a few hundred milliseconds for the one-to-three-sentence paragraphs Studio produces.
- Cues are capped at 42 characters (about two lines) and never shorter than half a second, so nothing flashes past unread.
SRT or VTT?
| Format | Use it for | Notes |
|---|---|---|
| SRT | Premiere, DaVinci, CapCut, YouTube upload, most LMSs | Comma decimal in timestamps; near-universal |
| VTT | HTML5 video, web players, some course platforms | Dot decimal; supports styling and positioning |
Getting them
- Produce in a Studio projectSingle clips from Text to Speech export audio only; subtitles come with Studio exports.
- Export a chapter or the bookSubtitles are on by default; both SRT and VTT are produced alongside the audio (Creator and above).
- Import into your editor or platformAttach as a caption track, or upload with the video. Edit any cue like normal text.
Two things to know
- Inline emotion tags such as [whispers] are stripped from the subtitle text — they direct the voice, they aren't said.
- On the Free plan, subtitle exports end with a short attribution cue after the last line; paid plans export clean files.
Translated editions
Each language edition of a translated project exports its own subtitles from its own timeline, so captions always match the narration in that language.
Frequently asked questions
Are the subtitles word-timed?
Sentence-timed within exact paragraph boundaries. For word-level karaoke-style captions, use the speech-to-text tool on the finished audio instead.
Can I edit the subtitles?
Yes — they're plain SRT/VTT files. Any editor or caption tool can change text and timing.
Do they work with YouTube?
Yes. Upload the SRT with the video or add it under Subtitles; YouTube accepts both formats.
Can I get subtitles for a single Text to Speech clip?
Put the clip in a Studio project (one paragraph is enough) and export; subtitles are a Studio export feature.