Glossary
Hertz (Hz) is the SI unit of frequency: one cycle per second. Named for Heinrich Hertz, who first conclusively demonstrated electromagnetic waves in 1887. In audio, Hz quantifies two things: the sample rate at which audio was digitized (samples per second) and the frequency content of the sound itself (cycles per second of the underlying waveform). Both matter for transcription, though sample rate matters more.
Audio frequency ranges
Human hearing: ~20 Hz to ~20,000 Hz (degrades with age, especially the high end).
Human speech: ~80 Hz (low male fundamental) to ~8,000 Hz (high-frequency consonant energy). The bulk of intelligibility lives between 300 Hz and 3,400 Hz — which is why telephony works at 8 kHz sample rate (capturing up to 4 kHz).
Music: ~30 Hz (low bass) to ~16,000 Hz (high cymbals, sibilance).
Below human hearing: infrasound. Above: ultrasound.
Hz in sample rates
Sample rate is also expressed in Hz. 8,000 Hz = 8 kHz = telephone quality. 16,000 Hz = Whisper's input rate. 44,100 Hz = CD quality. 48,000 Hz = professional video standard. The Nyquist–Shannon theorem: a sample rate of N Hz can represent frequencies up to N/2 Hz. So 16 kHz sample rate captures all frequencies up to 8 kHz — plenty for speech.
What Hz changes mean for transcription
Audio recorded at 44,100 Hz contains all human speech frequencies, then some. Audio at 8,000 Hz (phone calls) loses high-frequency consonants (especially /s/, /f/, /θ/), making them harder to distinguish. Whisper handles 8 kHz audio reasonably well — it's been trained on plenty of phone-quality data — but accuracy is 2-4 WER points worse than 16 kHz+ audio. If you're choosing a recording sample rate purely for transcription, 16 kHz mono is the cost-efficient sweet spot.
In practice
Your iPhone records voice memos at 44.1 kHz mono. Your phone-app conference recording is 8 kHz mono. Both are processed through the same Whisper pipeline at Transcript.you. The 44.1 kHz file may transcribe at 96% accuracy; the 8 kHz file at 93%. The difference is high-frequency consonant clarity, not the absolute Hz number.
Related terms
Further reading
Transcribe at any sample rate
Get started — freeLast updated: September 20, 2026
Frequently Asked Questions
The unit of frequency: cycles per second. Audio signals are described by their sample rate (e.g. 44,100 Hz) and frequency content. Human speech ranges roughly 80 Hz to 8,000 Hz.
Hertz (Hz) affects how you understand transcript quality, timing, compatibility, or the technology behind speech-to-text results.
Use the related workflow at /audio-to-text when you want to see how this glossary concept connects to an actual transcription task.
Browse the full transcription glossary or read the complete guide to AI transcription.