Why an auto-caption transcript reads badly
Automatic captions are written for a rolling window rather than for reading. Each cue restates the previous one and adds a few words, so the text scrolls smoothly on screen — and pasting a YouTube VTT straight into a document gives you the same phrases over and over.
Removing repeated lines is on by default. It drops a cue that simply repeats the one before it, and where a cue restates the previous line and extends it, it keeps only the longer version. Turn it off if a line is genuinely said twice in your file.
It will not try to stitch together cues that only partly overlap, such as "we are going" followed by "going to the shops". Working out where to join those means guessing, and a wrong guess deletes words from your transcript silently. Anything it cannot resolve is left in place for you to see.

