How captioning a video works here
The audio is extracted from your video and transcribed with word-level timing, then grouped into caption blocks. You get an editable transcript rather than a finished video, which keeps the captions separate from the picture.
Keeping them separate is usually what you want. A separate file can be corrected without re-encoding, replaced with a translated version, switched off by the viewer, and read by the platform — YouTube indexes uploaded caption text, so a captioned video becomes findable by what was said in it.




