Speech to Text: Transcribe Audio and Video in 30 Languages
Upload a recording or a video, get a transcript with a timestamp on every word, and export it as TXT, SRT or VTT. Auto-detects the language among 30; files up to 200 MB. Billed at 1 credit per character, from the same credits as Text to Speech.
3,000 free credits ≈ three minutes of speech · No credit card required · One-click Google sign-up
00:00:00,000 → 00:00:02,640Welcome back to the channel. Today we're breaking down
00:00:02,640 → 00:00:05,120the three habits that quietly drain your focus,
00:00:05,120 → 00:00:07,300and the one fix that took me a week to notice.
Illustrative SRT output: word timings grouped into cues of up to 42 characters. Real files carry your own text.
How it works
- Upload a fileDrop in audio (WAV, MP3, M4A, FLAC, OGG, Opus, AAC, WebM) or video (MP4, MOV, MKV, AVI, WMV, FLV, MPEG, M4V) up to 200 MB. Long recordings are processed in segments.
- TranscribeLet the language auto-detect or pick one of 30. The transcript comes back with a timestamp on every word.
- Edit and exportFix any misheard word in the app, then export TXT, SRT or VTT. Subtitle cues are grouped at up to 42 characters per line.
What makes LeapFun speech to text different
- Word-level timing, not sentence guesses
Every word carries its own start time, so captions line up with the speaker instead of drifting across a sentence. That is what makes the SRT/VTT export usable without re-timing.
- One credit pool for the whole workflow
Transcription bills 1 credit per character of transcript — no per-minute meter, no separate subscription. The same credits generate speech, clone a voice or narrate a book.
- Built for creators, not call centres
Video files go straight in; long recordings are segmented automatically; the export lands in your editor as captions or in Audiobook Studio as a script to re-narrate.
30 languages, auto-detected
English, Chinese (Mandarin), Cantonese, Japanese, Korean, Spanish, French, German, Russian, Italian, Portuguese, Arabic, Indonesian, Thai, Vietnamese, Turkish, Hindi, Malay, Dutch, Swedish, Danish, Finnish, Polish, Czech, Filipino, Persian, Greek, Romanian, Hungarian and Macedonian. Leave the language on auto and the transcriber identifies it; set it explicitly when a recording mixes accents or is mostly names.
Speech to Text vs. Studio subtitles
| Speech to Text | Audiobook Studio subtitles | |
|---|---|---|
| Input | Audio or video you already have | A script you narrate in LeapFun |
| How the timing is made | Recognition, word by word | From the narration timeline — no recognition, exact text |
| Best for | Interviews, recordings, videos from other tools, repurposing old content | Captions for audio generated here |
| Billing | 1 credit per transcript character | Included in the Studio export on every plan (free tier adds an attribution line) |
| Export | TXT, SRT, VTT | SRT, VTT alongside WAV/MP3 |
What's free, what's paid
| Plan | Price | Credits | What it unlocks |
|---|---|---|---|
| Free | $0 | 3,000 credits on sign-up + 1,000 each active month | Transcripts for personal and evaluation use · SRT/VTT carry an attribution line |
| Starter | $6 / mo | 40,000 credits / month | Clean subtitle files · commercial rights |
| Creator | $19 / mo | 200,000 credits / month | Roughly 20 hours of talks a month · unlimited Studio projects |
Transcription is billed at 1 credit per character of the transcript on every plan. Speech runs about 900 characters a minute, so 1,000 credits ≈ 1 minute of talk — the same rule of thumb as narration. Full tiers, yearly pricing and Pro / Team plans on the pricing page.
Frequently asked questions
Is LeapFun speech to text free?
Yes to start. Your 3,000 sign-up credits cover about three minutes of speech, and every active month adds 1,000 more. On the free tier the SRT/VTT export ends with a short attribution line; paid plans export clean files and include commercial rights.
How is transcription billed?
One credit per character of the transcript, on every plan, from the same credit pool as Text to Speech — there is no per-file charge. A 10-minute talk is roughly 1,500 words, about 9,000 characters, so about 9,000 credits.
Which file types and sizes are supported?
Audio: WAV, MP3, M4A, FLAC, OGG, Opus, AAC and WebM. Video: MP4, MOV, MKV, AVI, WMV, FLV, MPEG and M4V. Up to 200 MB per file; long recordings are transcribed in segments, so a 10-minute file is routine.
Do I get word-level timestamps?
Yes. Every word carries a start time, which is what makes the SRT/VTT export usable as captions; cues are grouped at up to 42 characters and a 0.8-second gap.
What is the difference between Speech to Text and Studio subtitles?
Speech to Text transcribes audio you upload — it listens. Studio subtitles are generated from the narration timeline of a LeapFun project, so the text is exact and no recognition is involved. Use Speech to Text for recordings you did not generate here.
Can I edit the transcript before exporting?
Yes. The app shows the text, the word timeline and the cue lines; correct any word, then export TXT, SRT or VTT.
Ready to transcribe?
Sign up in one click, get 3,000 credits, and turn your first recording into a timed transcript in minutes.
Transcribe a file