Closed caption generator

Create subtitles and text automatically with AI in 3 minutes

  • Transcribe audio/video to subtitles or text.
  • Translate audio/video into subtitles or text in any language.
  • 76 languages supported.
  • Save output as subtitles or plain text.
Trusted by 3000+ users.

Free: one 60-second file a day

Sign up for 3 a day

Mode:

Transcribe
Translate

Create a closed caption file for any video or audio in under three minutes. Export as SRT or VTT and upload it alongside your video, so viewers can switch captions on and off.

Closed captions, subtitles and open captions

The three get used interchangeably, but they are not the same thing. Closed captions (CC) are a separate file the viewer can turn on or off, and they transcribe the language actually being spoken. Subtitles usually mean a translation into a different language. Open captions are burned permanently into the video image and cannot be switched off.

This tool produces the first two: a separate, switchable caption file, either in the spoken language or translated. It does not burn captions into video, which is deliberate — a separate file stays editable, is read by search engines, and can be swapped without re-encoding.

Why closed captions matter beyond accessibility

Accessibility is the obvious reason, and in many contexts a legal one: WCAG 2.1 Level AA requires captions for prerecorded audio content, and public-sector and education bodies in a number of jurisdictions are held to it.

The less-discussed reason is reach. A large share of social video is watched with sound off, and caption files are machine-readable in a way that speech is not — platforms index caption text, which makes a captioned video findable by what was said in it.

Which format to upload where

All three come from the same job at no extra cost. Generate once, export whichever you need, or all of them.

  • YouTube, Vimeo and most LMS platforms accept SRT — use this unless you have a reason not to.
  • HTML5 <track> elements and web players require VTT (WebVTT).
  • TXT is a plain transcript with no timings, for show notes, blog posts or documentation.

Editing before you publish

Captions are shown in an editor with the audio, so you can correct recognition errors and adjust block boundaries before exporting. Names, technical terms and numbers are worth a specific check — they are where automatic recognition is weakest and where an error is most obvious to a viewer.

ByteDance logoBerkley university logoUniversity of Illinois logoAljazeera logo
Transcribe 76 languages to Subtitles or Plain Text

Transcribe

Transcribe Your Video/Audio to Subtitles Automatically.

76 Languages supported (See List)

Translate

Translate Video/Audio into Subtitles in Any Supported Language, Automatically.

76 Supported Languages (See List)

Translate 76 languages to Subtitles or Plain Text

Supports Noisy Audio

Translate or transcribe audios recorded in a noisy environment with great accuracy.

Multilingual audio

Have video where multiple speak different languages? No problem!

Support for Diverse Accents

Different people from different places speak the same language differently, our AI can work through it seamlessly.

Common questions

What is the difference between CC and subtitles?

Closed captions transcribe the spoken language of the video and are intended for viewers who cannot hear the audio. Subtitles are normally a translation for viewers who cannot understand the language. This tool does both: transcription keeps the original language, translation produces subtitles in another.

What format are the closed caption files?

SRT (SubRip) and VTT (WebVTT), plus a plain TXT transcript. SRT is accepted by YouTube, Vimeo and most learning platforms; VTT is the format HTML5 video players expect.

Can I upload the CC file straight to YouTube?

Yes. Export as SRT and add it under Subtitles in YouTube Studio. YouTube will use your file instead of its own automatic captions, which is generally more accurate and gives you control over the wording.

Do closed captions need to be word-for-word?

For accessibility compliance they should closely reflect what was said, including meaningful non-speech audio where relevant. The generated transcript is verbatim by default, and you can edit it before export if you need to tidy filler words.