01 It reads captions instead of transcribing
YouTube already publishes a caption track for most videos — creator-uploaded when there is one, auto-generated otherwise. Reading that track takes seconds and costs nothing, where transcribing the audio from scratch means downloading the video and running a speech model over it.
02 Three formats, one click
Take a plain .txt transcript for reading, summarising or pasting into a document, or timed .srt / .vtt cues for an editor, a re-upload or a translation pass.
03 It is honest when it cannot help
If a video publishes no caption track at all, you are told so and pointed at the transcription tool, rather than being quietly charged for a speech-to-text run you did not ask for.