Glossary
Prosody is the umbrella term for the rhythmic, intonational, and stress patterns of spoken language. It's how something is said, separate from what is said. Pitch rises at the end of a yes-no question; emphasis falls on the contrastive word; speech rate slows for important points. Modern ASR mostly transcribes the words and discards the prosody, but prosody carries 30-40% of the meaning of conversational speech.
What prosody encodes
Three big things. (1) Sentence type: rising pitch = question, falling pitch = statement. (2) Focus and emphasis: "I didn't say he stole the money" vs "I didn't say he stole the money" — same words, opposite meanings, prosody disambiguates. (3) Emotion and attitude: sarcasm, doubt, certainty, excitement, boredom. Prosody is a major channel for these signals.
Prosody and ASR
Standard transcription discards prosody — the output is just words. Some systems annotate it: speech-and-language pathology research uses ToBI (Tones and Break Indices) to mark prosodic features in transcripts. For commercial ASR, prosody is mostly used implicitly by the language model to disambiguate (a question intonation biases toward question-mark output). Direct prosodic transcription is a research niche.
The frontier: prosody for emotion and intent
Audio LLMs (Whisper-fine-tuned, OpenAI Realtime, GPT-4o-audio) now consume audio directly without first transcribing to text — preserving prosody. This unlocks emotion detection ("the user sounds frustrated"), intent detection ("is this a question or a statement?"), and sarcasm detection. Expect prosody-aware transcription to become a standard feature in 2026-2027.
In practice
Recording: "Oh, great." Transcribed text alone is ambiguous — it could be enthusiasm or sarcasm. Prosody disambiguates: rising pitch + bright tone = enthusiasm; flat pitch + slow drawl + emphasis on "great" = sarcasm. Standard transcripts can't capture this; emotion-aware audio models can.
Related terms
Further reading
Try AI transcription
Get started — freeLast updated: September 06, 2026
Frequently Asked Questions
The patterns of stress, intonation, and rhythm in spoken language. Modern ASR mostly ignores prosody for word-level transcription, but it's the next frontier for emotion and intent detection.
Prosody affects how you understand transcript quality, timing, compatibility, or the technology behind speech-to-text results.
Use the related workflow at /ai-transcription when you want to see how this glossary concept connects to an actual transcription task.
Browse the full transcription glossary or read the complete guide to AI transcription.