Captions vs Subtitles vs Transcripts: What’s the Difference?
Understand how captions, subtitles, and transcripts differ in timing, sound information, accessibility, and production use.
The terms captions, subtitles, and transcripts overlap in everyday speech, but the deliverables are not interchangeable. A transcript captures spoken content in reading order. Subtitles synchronize dialogue or translation to video. Captions also represent meaningful non-speech audio for viewers who may not hear it.
Three deliverables from one scene
Video audio: a door slams, Maya whispers “He’s here,” and tense music begins.
Transcript:
Maya: He's here.
A verbatim transcript might add [door slams], but many editorial transcripts focus on speech.
Translation subtitle:
18
00:01:12,000 --> 00:01:14,000
Il est là.
It translates dialogue for an audience that can hear the original sound.
English closed caption:
18
00:01:11,500 --> 00:01:14,500
[door slams]
MAYA: He's here.
[tense music]
It includes sound cues and speaker identification needed to understand the scene without audio.
Comparison
| Feature | Transcript | Subtitle | Caption |
|---|---|---|---|
| Timed to video | Optional | Yes | Yes |
| Dialogue | Yes | Yes | Yes |
| Non-speech audio | Sometimes | Usually limited | Yes when meaningful |
| Speaker identification | Editorial choice | When needed | When not visually clear |
| Typical use | Search, notes, articles | Language access | Deaf/hard-of-hearing access |
“Open captions” are burned into the image and cannot be turned off. “Closed captions” are a selectable track. A selectable track can be delivered as SRT, VTT, TTML, or a platform-specific format; the concept and file extension are separate.
Why the distinction changes editing
Removing [music] from captions may remove story information. Adding every background sound to translation subtitles can overload viewers who already hear it. Turning a transcript into captions requires segmentation, timing, speaker labels, and sound descriptions—not merely adding a timestamp every few sentences.
YouTube terminology
YouTube’s interface often groups subtitle and caption tracks under “Subtitles.” Inspect the actual track and audience requirement. An automatically generated speech track may not include meaningful sound effects, so it is not a complete accessibility caption file even if the UI calls it captions.
Choose the deliverable before production
Ask who will use it, whether audio is available, which language is needed, whether the text must be searchable, and what platform will receive it. Then specify style rules for sound cues, speaker labels, music, profanity, and verbatim speech.
To build timed text from a video with no track, follow How to Create Subtitles for a YouTube Video Without Captions. For extraction from an existing track, use How to Extract Subtitles from a YouTube Video. The complete YouTube subtitle workflow connects both sources to cleanup, translation, QA, and dubbing.