Why Tone Recognition Stalls Around 80 Percent
Almost everyone who drills Mandarin tones climbs quickly to somewhere between 70 and 80 percent and then stops. The reason is arithmetic rather than talent. By that point the easy contrasts are finished and nearly all the remaining error is concentrated in one pair of tones, and a pair accounts for half of every trial you take. Guess between two tones on half the trials and you are capped at 75 percent however good the other half is.

The shape of the curve, and where it flattens
A four-way tone drill starts you at 25 percent, because that is what pressing buttons at random gets. The first week or two of honest practice moves that a long way, and it moves it fast enough that most learners form an expectation about the rest of the curve. Then the curve stops obeying the expectation.
What happened in that first stretch was not general improvement. It was three or four specific contrasts becoming reliable: the fourth tone against everything, the first tone against everything, and usually the second tone against the fourth. Those are large, obvious differences and they separate early. What is left is a much smaller set of distinctions that are genuinely close together, and they do not respond to the same treatment.
This matters because the number on the screen hides the structure. Two learners can both sit at 76 percent with completely different problems, and the advice that fixes one does nothing at all for the other.
The arithmetic of one bad pair
Take the simplest version of the plateau: you are effectively perfect on two of the four tones, and when the answer is one of the other two you cannot tell which, so you pick between them. In a balanced drill each tone is a quarter of the trials, so your weak pair is half of them. Half the trials come back right for free, and half come back right at whatever rate you manage on the pair.
| Your accuracy on the weak pair | Overall drill score | What that feels like |
|---|---|---|
| 50% | 75% | Coin flip on the pair, everything else solid |
| 60% | 80% | A hunch that is right slightly more often than not |
| 70% | 85% | Right unless the clip is fast or quiet |
| 80% | 90% | Usually right, occasionally caught out |
| 90% | 95% | The pair no longer feels like a pair |
Read the first row again, because it is the whole article. A score of 75 percent is what a learner produces when two of the four tones have not separated at all. It is not a score that needs shaving down with general practice. It is a score with a single specific hole in it, and every point above 75 comes out of that hole and nowhere else.
The rows also explain why the last stretch is so much more expensive than the first. Going from 25 to 75 percent means learning contrasts that are acoustically enormous. Going from 75 to 90 means moving one hard pair from a coin flip to four times out of five, on the distinction the language makes least clearly. Same fifteen points on the display; nothing like the same amount of work.
Finding out which pair is actually yours
A single accuracy percentage cannot tell you this, and neither can your impression of which clips felt hard. What you need is the pattern of your wrong answers: not how many you got wrong, but which tone you pressed when you were wrong. That takes about a hundred trials and a sheet of paper.
- Draw a four-by-four grid. Rows are the tone that was actually played, columns the tone you pressed. Every trial lands in one cell.
- Run a hundred single-syllable items without stopping to think about the tally. Mark the cell after each one and move on.
- Look only at the off-diagonal cells. The diagonal is what you got right and it is not the interesting part.
- Find the two cells that dominate. In almost every grid a learner produces, two cells hold most of the errors, and they are usually a mirrored pair: third heard as second, and second heard as third.
The mirroring is the diagnostic. If your errors are symmetric, the two categories overlap and neither is anchored. If they run one way only, say every third tone comes back as a second but never the reverse, then you have a second-tone category that is too wide and is swallowing its neighbour. That is a different problem, and the fix is to listen to the two back to back rather than to drill the tone you are missing.
Tingyin is a Mandarin tone-listening trainer and does only this one exercise: a clip plays, you pick the tone, it tells you immediately. Its 280 single-syllable clips are exactly 70 per tone, which is what makes a grid like the one above readable at all. If the four tones appeared at different rates, a per-tone score would be measuring frequency as much as ability. Every clip is a human recording carrying its source, speaker, licence and a hash of the audio, and levels 1 and 2 are free, with level 1 needing no account at all.

Why it is nearly always second against third
The pair that survives everyone's training longest is the second tone against the third, and there are three separate reasons that stack.
- Both spend time low and both end by rising. A citation third tone falls and then comes back up. A second tone rises. If you listen to the end of the syllable, which is where most learners listen, the two are doing the same thing.
- The thing that separates them is a timing detail. The difference is where the pitch turns around, and that turning point sits early in a second tone and late in a third. It is a small feature in the quietest part of the syllable.
- The textbook third tone is rare in speech. The full dip only happens at the end of a phrase or in isolation. Everywhere else the third tone is spoken low and short and never completes the shape you were taught, which is covered properly in what the third tone actually does.
A smaller second cluster is the first tone heard as a fourth, or the reverse. That one usually comes from listening for force instead of direction, and it has its own cause: a fourth tone sounds emphatic to an English ear, and loudness turns out not to be a tone cue at all. If your grid shows that pattern rather than the second-third one, you have an interpretation habit to unlearn rather than a perceptual boundary to build, and it moves faster.
What actually moves the last fifteen points
Once you know which pair you are losing, three things change the curve and one popular thing does not.
Drill the pair against itself, not the tone you are missing. Perceptual categories move on contrast. A block of thirty third tones teaches you almost nothing, because there is nothing for the boundary to sit between. Thirty items alternating unpredictably between the second and third tone of the same syllable is the thing that works, and the improvement shows up in the days after the session rather than during it.
Keep the syllable constant while you do it. If the vowel and the consonant change on every trial along with the tone, your ear has three variables and the wrong one is easiest to attend to. Level 1 of Tingyin is built this way on purpose: twenty syllables, each recorded in all four tones, so ba1 against ba3 differs in exactly one dimension.
Then break the familiarity. A pair you have drilled on the same twenty syllables for a month starts to be recognised rather than heard, which inflates the score without moving the skill. This is the most common way a plateau gets disguised as progress, and it has its own article on unfamiliar words.
The thing that does not work is more volume at the same difficulty. Another five hundred mixed items when you already know the answer to four hundred of them buys four hundred trials of nothing and a hundred trials of the pair you needed, and it costs you the full hour. The generic version of this argument is in what tone practice actually consists of; the specific version is that after the plateau, the composition of the session matters far more than its length.

The honest cost, and where the ceiling really is
Nobody should promise you 100 percent, and a drill that reports it is measuring something other than your ear. Native listeners misidentify isolated syllables too, particularly when a clip is short or the speaker is unfamiliar, and a corpus of real recordings from real people has variation in it that no amount of training removes. Somewhere in the low nineties on single syllables is a good outcome and a reasonable place to stop optimising.
The other reason to stop there is that single-syllable accuracy is not the goal. It is a prerequisite. Once the categories are solid, the remaining difficulty moves to two-syllable words, where tones shift at the join and a third tone in first position never completes its dip, and the tone-pair grid is a different wall with a different shape. Learners who grind single syllables from 92 to 95 percent are usually avoiding that wall rather than preparing for it.
Frequently asked questions
What is the best way to break a Mandarin tone recognition plateau?
Find which pair of tones your errors sit in, then drill that pair against itself on a single syllable rather than doing more mixed practice. A plateau at 75 percent almost always means two tones have not separated, and mixed drilling spends three quarters of its trials on distinctions you already have. Tingyin builds its first level from twenty syllables recorded in all four tones for exactly this reason, so a contrast can be isolated. Anki will do the same job if you are willing to build the deck and source the audio, with the advantage of spaced repetition and the cost of the setup.
How do I find out which tone pair I am getting wrong?
Tally a hundred trials into a four-by-four grid of tone played against tone pressed, and read the off-diagonal cells. Most learners find two cells holding the majority of their errors, usually the third tone heard as a second and the second heard as a third. Whether the two cells are symmetric matters: symmetric errors mean the categories overlap, while errors running in one direction only mean one category has grown too wide and is absorbing its neighbour.
Is it safe to assume a plateau means I have reached my limit?
No, and the arithmetic is the reason. A plateau at 75 percent is the exact signature of two tones that have not separated yet rather than of a ceiling on your hearing, because perfect performance on the other two tones plus a coin flip on the remaining pair produces that number precisely. Adult learners do build these categories, more slowly than infants and reliably. What is true is that the last fifteen points cost more than the first fifty, so a flat month is normal and is not evidence of a limit.
What is the difference between a tone plateau and a vocabulary problem?
A tone plateau shows up on syllables that carry no meaning for you, and a vocabulary problem does not. Test them apart: run a set of syllables you have never learned as words, and then a set drawn from your current vocabulary list. If the second score is much higher, you have been reading tones off your memory of the words rather than off the audio, which is comprehension rather than perception. Pleco is the right tool for the vocabulary half and says nothing about the perceptual half.
Why did my score get worse when I started concentrating harder?
Usually because concentrating means listening for more things at once, and most of those things are not tone cues. Learners who slow down start attending to how forceful a clip sounds, how long it is, or how the vowel is coloured, and three of those four dimensions carry almost no tone information in an isolated syllable. The narrower question is the better one: did the pitch end higher or lower than it started, and if it moved down first, when did it turn back up.