Whispered Mandarin Still Has Tones
Whispering removes vocal fold vibration, which removes pitch, which should remove tone entirely. It does not. Listeners identify tones in whispered Mandarin well above chance, though well below normal speech, and they do it from what is left over: how long the syllable is, how its energy is shaped across that time, and small shifts in vowel quality that travel with the pitch a speaker intended. That leftover set is a good description of what a tone is made of.

What whispering actually removes
Normal speech is powered by the vocal folds opening and closing in a cycle, a hundred and something times a second for most adults. That repetition rate is the fundamental frequency, and it is the physical thing your ear turns into pitch. A Mandarin tone is a pattern traced by that rate over the course of a syllable: held level, rising, dipping and returning, falling.
In a whisper the folds do not vibrate. Air is pushed through a narrowed glottis and the result is turbulence, which is aperiodic by nature. There is no repetition rate to measure, so there is no fundamental frequency, so there is no pitch and no contour. On paper the tonal system of the language has just been switched off.
This is not a partial reduction, the way a phone call or a small speaker degrades things. Those keep the harmonic structure and lose parts of it, which is why tones survive a phone line almost intact. A whisper has no harmonic structure to keep. Everything the previous paragraph described as carrying tone is genuinely, completely gone.
And people still get it
Whispered Mandarin is a real register, used the way whispering is used in any language, and native speakers understand each other in it. When researchers have tested this directly, by asking listeners to identify the tone of whispered syllables presented in isolation, the pattern that comes back is consistent: accuracy lands clearly above the 25 percent a four-way guess would give, and clearly below what the same syllables produce when spoken normally.
Both halves of that sentence matter. The gap from chance says information about tone is present in a signal that contains no pitch. The gap from normal speech says pitch was doing most of the work and the leftovers are a poor substitute. Neither the strong claim nor the dismissive one is right, and the interesting question is what specifically is in the leftovers.
The three things left over
Tingyin is a Mandarin tone-listening trainer, and every clip it ships is a human recording carrying its source, speaker, licence and a hash of the audio. Those 637 clips are all normal phonated speech; there are no whispers in the corpus and it cannot settle what a whisper preserves. What it can do is show how large the non-pitch dimensions are in ordinary speech, which is the raw material a whisper has to work with.
| Cue | Survives a whisper? | How much it separates the four tones |
|---|---|---|
| Pitch contour | No | Almost all of it, in normal speech |
| Duration | Yes | Separates the third tone; barely separates the other three |
| Shape of energy over time | Yes | A real difference in average level, up to about 4 dB |
| Vowel quality and formants | Yes | Small, and not measured in this corpus |
| Peak loudness | Yes | Nothing at all |
Duration is the strongest of the three and it is not strong. Across the 280 single-syllable clips, 70 per tone, a third tone averages 1.330 seconds while the other three sit at 1.052, 1.083 and 1.012. The third tone is genuinely long and the rest are essentially the same length as each other, which means duration hands a whispering listener one of the four tones and leaves a three-way problem behind. The full length breakdown makes the same point at more depth.
The energy shape is subtler and it needs care. The peak level of all four tones in this corpus falls within 0.13 dB, so loudness in the sense of how loud a syllable gets tells you nothing. The average level over the whole clip does differ, spanning about four decibels from the first tone to the third, and it differs because the four tones distribute their energy across time differently. A fourth tone reaches full level early and spends the rest of its short life coming down. A first tone gets there and stays. That envelope is a property of how the syllable is articulated rather than of how the folds are vibrating, and articulation is exactly what a whisper keeps.

The part that is genuinely strange
The third leftover is the one that changes how you think about tone. Formants are the resonances of the vocal tract, the frequency bands that a particular tongue and jaw position amplifies, and they are what makes an a sound like an a rather than an i. They are determined by the shape of the cavity above the larynx, not by the vibration below it, so a whisper preserves them.
The finding, reported in work on whispered tone across several tone languages, is that these resonances are not independent of the tone a speaker is producing. Raising pitch involves tensing and lifting the larynx, and moving the larynx changes the length of the tube above it, which shifts the resonances slightly. So a syllable intended as a high tone has subtly different formant values from the same syllable intended as a low tone, whether or not the vocal folds are vibrating at all.
That is a mechanical side effect, not a designed signal, and it is small. But it means the tone leaves a trace in the timbre of the vowel and not only in its pitch, and it means a whispering speaker can produce that trace without doing anything special. Listeners appear to pick some of it up.
None of this is measured in Tingyin's corpus. The manifest records duration and two loudness figures per file and does not record formants, and nothing anywhere in the project analyses the pitch of a clip or verifies that a recording carries the tone it is labelled with. The audio is trusted because of where it came from and what it is: the checks confirm the bytes, the licence and the hash. Take the formant claim as a report from the literature rather than as something this repository has verified.
What it teaches about what a tone is
The usual mental model of a Mandarin tone is a melody laid on top of a syllable, as though the syllable were a note and the tone a way of singing it. Whispering breaks that model in a useful way. Take the melody away and something recognisable remains, which means a tone was never a melody laid on top. It is a coordinated gesture of the whole vocal apparatus, and pitch is its loudest consequence rather than its definition.
That reframing has a practical edge for a learner. The gesture bundles several things together, and every one of them is a hint:
- How long the syllable runs. The third tone dips and returns and that takes time, in a whisper as much as in speech.
- Where in the syllable the energy peaks. Early and falling away is a fourth tone shape; sustained is a first tone shape.
- How the vowel is coloured. A trace, not a cue you can deliberately train, but it is there.
In normal speech these are redundant with the pitch contour, which is why nobody teaches them and why they should not become your primary strategy. Redundancy is what makes speech robust, and a whisper is simply the condition that strips the signal back to the redundant layer and lets you see what was underneath.

What this does not mean
Three conclusions people draw from this fact are wrong, and they are worth naming.
It does not mean tones are optional. Whispered tone identification is above chance and a long way below normal, and whispered conversation leans heavily on context, familiar vocabulary and lipreading to make up the difference. A whisperer is not communicating tone efficiently; they are communicating badly and being rescued by everything around the syllable. That rescue is exactly the mechanism described in why tone accuracy drops on words you do not know.
It does not mean you should practise on whispers. A drill exists to build the pitch categories, and the fastest way to build them is a clean signal with unambiguous feedback. Training on a degraded version of the cue teaches you to lean on the substitutes, which are weaker in every case. Tingyin serves full-bandwidth human recordings for this reason, and levels 1 and 2 are free, with level 1 needing no account at all.
It does not mean tone is really about duration and loudness. The leftovers separate one tone cleanly and the other three poorly, which is what you would expect from a set of side effects. If they were the real signal, whispered accuracy would be near normal, and it is not. Pitch is doing the work; the leftovers are what falls out of doing it.
Frequently asked questions
Can you speak Mandarin while whispering?
Yes, and native speakers do it constantly, in libraries and cinemas and next to sleeping children. Tone identification from whispered syllables in isolation lands above chance and well below normal speech, so a whispered conversation relies more than usual on context, familiar words and watching the speaker's mouth. It works because conversation has redundancy, not because the tones came through cleanly.
How do tones survive without pitch in a whisper?
Through three side effects of producing a tone that do not depend on the vocal folds vibrating. The syllable's duration is one, and it separates the third tone from the rest reliably. The way energy is distributed across the syllable is another, since a fourth tone peaks early and falls away while a first tone sustains. The third is a small shift in vowel resonance, because raising pitch moves the larynx and moving the larynx changes the shape of the tube above it.
What is the difference between whispering and speaking quietly?
Speaking quietly still uses vocal fold vibration, so it still has a fundamental frequency and a full tone contour, just at a lower volume. Whispering has no vibration at all and therefore no pitch of any kind. That is why quiet speech carries tone perfectly well and whispered speech does not, and it is also why turning down the volume on a drill does not simulate the whispered condition.
Is it safe to practise Mandarin tones by whispering to myself?
Not for tone work, no. Whispering removes the one dimension you are trying to train and leaves the weakest substitutes in place, so it builds the habit of listening to duration and emphasis rather than to pitch direction. If you need silent practice, listening on headphones is the better option, and Tingyin's level 1 runs in a browser without an account. Whispering is fine for rehearsing vocabulary or word order, where the pitch was not the point.
Why does whispered Mandarin sound flat rather than wrong?
Because the tonal information is reduced rather than replaced. A tone spoken with the wrong pitch contour is actively misleading and points at a different word, while a whispered tone points weakly at the right one. Listeners experience the first as an error and the second as an absence, which is why whispering feels like a loss of expressiveness rather than like speaking a different language.