Upload a file, pick a language, and get back a caption track with word-level timings you can edit before you export it. Most jobs finish in under three minutes regardless of how long the recording is.
What the AI actually does
Captions are generated by ElevenLabs Scribe v2, a speech recognition model trained for transcription rather than for conversation. It returns a timestamp for every individual word, not just for each caption block, which is what makes the timings line up with speech instead of drifting a second or two out over a long recording.
That word-level output is also why re-timing is cheap. When you split or merge a caption in the editor, the new block inherits real timings from the words inside it — nothing is interpolated or guessed.
Where automatic captions usually go wrong
The failure modes worth knowing about are accents, overlapping speech, background music and code-switching — a speaker moving between two languages mid-sentence. Scribe v2 handles all four noticeably better than the free auto-caption features built into video platforms, but none of them are solved problems.
This is why the editor exists and why captions are never exported automatically. Read the transcript before you publish it, especially for proper nouns, brand names and numbers, which are where recognition errors concentrate and where they are most visible to a viewer.
Captions in one language, or translated into another
Transcription keeps the spoken language. Translation takes the same audio and produces captions in any of the 76 supported languages, so a Hindi interview can become English captions, or an English webinar can become Spanish ones, in a single pass.
Translation runs on the transcript rather than on the audio, so timings are preserved exactly — the translated caption appears at the moment the original words were spoken.
What it costs
You can try it without paying: one file up to 60 seconds per day without an account, or 3 a day with a free one.
Beyond that, credits cost $10 for 100, one credit per minute of audio or video, rounded up to the next whole minute per file. A 12-minute podcast episode costs 12 credits. There is no subscription, and credits expire 365 days after purchase.
Transcribe
Transcribe Your Video/Audio to Subtitles Automatically.
Accuracy depends mostly on audio quality. Clear single-speaker recordings in a well-supported language are typically accurate enough to publish after a quick read-through. Heavy background noise, crosstalk and strong regional accents all reduce it, which is why every transcript is editable before export.
Can I generate captions without a video file?
Yes. Audio-only files work exactly the same way — mp3, wav, m4a and similar. If you only need the words rather than a timed caption track, export as TXT instead of SRT or VTT.
Does it add a watermark?
No. The output is a caption file — SRT, VTT or TXT — not a rendered video, so there is nothing to watermark. The file is yours to use however you like.
Is there a free trial?
Yes. You can caption one file of up to 60 seconds per day without creating an account, and 3 files a day with a free account. Longer files and higher volume need credits, which cost $10 for 100 minutes and expire 365 days after purchase.