Transcript vs Subtitles vs Captions: What Is the Difference?
By the Speakmi team · Updated October 2, 2026 · 2 min read
People use these words interchangeably, but they are different tools. Choosing the right one saves time and avoids confusion for your audience.
Transcript
A transcript is the spoken content as plain text, usually without timing. It is best for reading, searching, quoting, and publishing next to a podcast. The common file format is TXT or a word processor document.
Subtitles
Subtitles assume the viewer can hear the audio but does not understand the language. They show the dialogue, often translated, in sync with the video. They usually leave out sound effects.
Captions
Captions are written for viewers who cannot hear the audio. They include the dialogue and also important sounds such as music, laughter, or a door slamming, plus speaker names when needed. Closed captions can be switched on and off, while open captions are burned into the picture.
File formats
- SRT: the simplest and most widely supported timed text format.
- VTT: the standard for web video and supports basic styling.
- TXT: plain text for transcripts.
Which should you choose?
- Publishing a podcast or interview: a transcript.
- Reaching viewers who speak another language: subtitles.
- Making a video accessible: captions.
The voice.speakmi transcriber produces the text and timing for all three. You then decide how to present it, and add sound descriptions yourself if you need true captions.