Glossary
Transcription and subtitling start with the same audio and the same ASR output, but produce very different deliverables. Transcription is a written document of what was said. Subtitling is timed on-screen text optimized for reading speed and screen real estate. Knowing which you need affects format, length, and editing decisions.
What transcription gives you
A document — DOCX, TXT, PDF, Markdown — with the spoken words written out, often with speaker labels and segment timestamps at paragraph boundaries. Optimized for reading top-to-bottom: full sentences, no line-break constraints, complete punctuation, often light cleanup of disfluencies ("um", "uh", "you know"). Typical use: research notes, articles, court reporting, blog drafting.
What subtitling gives you
A time-coded file (.srt or .vtt) with the spoken words split into short cues, each optimized for on-screen reading. Industry conventions: 32-40 characters per line, max 2 lines on screen at a time, 1-7 seconds per cue, reading speed ≤17 chars/second. Optimized for the viewer who's also watching the video: short, punchy, never blocking the action.
Choose based on what your audience does with it
Will they read the output without watching the video? Use transcription. Will they watch the video and read along? Use subtitling. Need both? Generate the transcript and a subtitle file in parallel — we export DOCX, SRT, VTT, and TXT in one upload. The underlying ASR output is the same; only the formatting differs.
In practice
An hour-long podcast interview: transcription output is a 12,000-word DOCX with speaker labels and timestamps every 30 seconds — readable as an article. Subtitle output is a 700-cue SRT file averaging ~15 words per cue, suitable for embedding in YouTube or Vimeo so viewers can read along.
Related terms
Generate subtitles
Get started — freeLast updated: September 21, 2026
Frequently Asked Questions
Transcription produces a plain text file with everything spoken. Subtitling produces time-coded short lines optimized for on-screen reading (typically 32-40 chars per line, max 2 lines on screen).
Transcription vs subtitling affects how you understand transcript quality, timing, compatibility, or the technology behind speech-to-text results.
Use the related workflow at /audio-to-srt when you want to see how this glossary concept connects to an actual transcription task.
Browse the full transcription glossary or read the complete guide to AI transcription.