Mandarin Tones Do Not Flatten in Fast Speech
Mandarin tones do not flatten when a speaker speeds up. Across three speaking rates, the pitch excursion of the second and third tones showed no effect of rate, and neither did the point where the pitch turns. What changes is the tone before it, and that is exactly what an isolated clip, in Tingyin or in Pleco, takes away.

What speeding up actually does to a tone
Almost every learner arrives at the same theory after their first real conversation: native speakers go too fast for the tones to survive, so the contours must be getting squashed. It is a reasonable theory. It is also the one thing in this article that has been measured directly and found not to be true.
Joan Sereno, Hyunjung Lee and Allard Jongman recorded twelve Mandarin speakers producing fifteen monosyllables with all four tones inside a sentence frame, at slow, normal and fast rates, and reported the results at ICPhS 2015. They measured vowel duration, consonant duration, the turning point where a rising or dipping tone changes direction, and the size of the pitch change itself. Two of those four moved with speaking rate. Two did not.
| Measure | Fast against normal | Effect of rate |
|---|---|---|
| Vowel duration | 19% shorter | Significant, and the same across all four tones |
| Consonant duration | 10% shorter | Significant, regardless of tone |
| Turning point location | No reliable change | None: F(2,3) = 0.747, p = 0.545 |
| Pitch excursion (change in F0) | No reliable change | None: F(2,3) = 0.415, p = 0.693 |
The two tones that carry the whole burden of this stayed apart at every rate. Turning point averaged 174 ms for the third tone against 58 ms for the second, and the pitch excursion averaged 37 Hz against 6 Hz, with no interaction between rate and tone anywhere in the data. The authors put it plainly: even at fast rates, speakers were able to maintain the changes in fundamental frequency that distinguish the second tone from the third.
So the syllable gets shorter and the pitch does not. A fast talker is giving you the same contour in less time, which is a real difficulty, but it is a different one from the theory most learners are carrying around.
The thing that does change is the syllable before it
In the same study, the preceding tone did move the turning point. With a first tone in front of it, the second tone's turning point ran 20 ms longer than with a fourth tone in front. For the third tone the same comparison moved only 6 ms. Rate did nothing; the neighbour did something, and it did different amounts of it to different tones.
Yi Xu measured how far that reaches, in a 1997 paper in the Journal of Phonetics that is still the baseline for the effect. At the boundary between two syllables, the starting pitch of the second tone varies by as much as 55 Hz on average depending on which tone came before it. The difference is still around 35 Hz at the onset of the following vowel, still around 17 Hz at its midpoint, and still measurable at the offset. His summary of what that does to a contour is worth reading slowly: the shape of a following tone can sometimes be distorted beyond easy recognition by the eye without taking the preceding tone into account.
The direction is asymmetric, which is the part that makes it hard to intuit. Carry-over effects are assimilatory, meaning the start of a tone is dragged toward wherever the previous tone ended. Anticipatory effects run the other way and are dissimilatory: a low onset coming up raises the peak of the tone before it rather than lowering it. The backward pull is large; the forward one is small.

Isolation is the harder case, not the easier one
Here is the finding that should change how you think about a drill. In the ICPhS data, all of it collected in sentence context, there was no overlap at all between the turning point of a second tone spoken at the slowest rate and a third tone spoken at the fastest. The two categories stayed cleanly separated across the entire range of speeds tested.
In isolated words, measured in an earlier study by Sereno and Lee, they do overlap. Strip the sentence away and the second and third tones start colliding on the very cue that is supposed to tell them apart. Context is not the thing that makes tones hard to hear. Context is carrying part of the signal, and a clip is the condition in which that help has been removed.
This inverts the usual complaint. A learner who scores well on clips and then misses speech is not failing at a harder version of the same task. They are failing at a different task, one that adds transitions between tones and takes away the clean edges a single syllable has. Drilling tone pairs rather than single tones is the first bridge across that, because a pair is the smallest unit that contains a transition at all.
What a clip drill is genuinely for
Tingyin is a Mandarin tone-listening trainer with one exercise: a clip plays and you pick the tone. Every clip is a human recording, and the shipped manifest carries the source, speaker, licence and sha256 of each of the 637 of them, 622 under CC BY-SA 4.0 and 15 under CC0. Level 1 is free with no account, and level 2 is free once you sign in.
Given what the measurements above say, the case for a clip is not that it simulates conversation. It is that it isolates the one decision that has to become automatic before anything else can, in the condition where the two easiest tones to confuse are hardest to tell apart. If you can separate a second tone from a third with no sentence around it, you have more margin than a conversation will ever ask for on that particular cue.
What that does not buy you is the transitions. 280 of Tingyin's clips are single syllables and 357 are two syllables or more, and it is only the second group that contains a boundary of the kind Xu measured. The single-syllable clips train a category; the multi-syllable ones are where you start hearing one tone leaning into the next.
Closing the rest of the gap
The honest answer is that clip drills stop paying somewhere around the point where you are reliably right on isolated syllables, and everything after that is connected speech. Three things move you across:
- Two-syllable material, deliberately. A pair contains one transition, which is enough to start hearing carry-over without drowning in it. This is also where tone sandhi and the neutral tone begin to matter, because both are changes that only exist between syllables.
- Audio you already understand. Re-listening to a sentence whose meaning you know removes the lexical work and leaves the tonal work, which is the only condition in which you can actually attend to a transition.
- Slow speech last, not first. Slowing a recording lengthens the syllable without restoring the contour information you were missing, because the contour was never the part that speed removed.

Frequently asked questions
What is the best way to practise Mandarin tones in fast speech?
Work on two-syllable material you already understand, rather than on faster versions of single syllables. Tingyin has 357 clips of two syllables or more out of 637, which is where a transition between tones actually exists. Du Chinese and other graded-reader apps cover the longer end of this better than any tone drill does, because the point is following a sentence rather than labelling a syllable.
How do I train my ear for the second and third tones at speed?
Listen for where the pitch turns rather than how fast the syllable goes past, because measurement shows the turning point does not move with speaking rate. In the ICPhS 2015 data it averaged 174 ms for the third tone against 58 ms for the second at every speed tested. If the pair is failing for you even in isolation, what the third tone actually does is the thing to fix first.
Is it safe to slow down a recording to hear the tones?
It is harmless and it is usually the wrong tool. Slowing playback lengthens the syllable, and the syllable's length was never the missing information: pitch excursion showed no rate effect at all in the measurements, so there is no extra contour hiding in the slow version. Re-listening at normal speed to something you have already understood teaches you more, because the difficulty was context rather than duration.
What is the difference between tone sandhi and tonal coarticulation?
Sandhi is a categorical rule: two third tones in a row, and the first is produced as a rising tone. Coarticulation is gradient and has no rule to learn, just a pull, with the start of one tone assimilated to wherever the previous one ended. Xu measured that pull at up to 55 Hz at the syllable boundary, which is large enough to reshape a contour without ever changing which tone was intended.
Why can I ace the drills and still not follow a conversation?
Because those are two different tasks and the drill is genuinely the cleaner of the two. Tingyin gives you one syllable with nothing around it; a conversation gives you transitions, a neutral tone here and there, and no time to decide. The useful reframe is that the drill was never a simulation of speech, and being good at it means the tone categories are solid, which is a prerequisite rather than the finish line.