Tone Accuracy Drops on Words You Do Not Know
If your tones are reliable on vocabulary you know and unreliable on syllables you have never met, your ear is fine and your test is broken. Knowing the word supplies the answer before the audio finishes, and a score built from familiar material measures recall wearing a listening costume. Unfamiliar syllables are the only honest measurement, and they come back lower for almost everyone.

The split most learners eventually notice
It usually shows up somewhere around the six-month mark. Drills from your vocabulary list come back in the nineties. Then someone says a name, or a place, or a word you have not met, and the tone is simply not there. Not wrong, not half-heard, absent. The syllable arrives as a bare sound with no tone attached to it.
The instinctive reading of that experience is that unfamiliar words are somehow harder to hear, that they are spoken faster or less clearly. Occasionally true and mostly not. The same speaker, the same recording conditions, the same length of syllable can produce a confident answer when you know the word and nothing at all when you do not. The audio did not change. What changed is that on the familiar word you were not really listening.
What knowing the word actually hands you
A known word narrows the field before the sound has finished arriving, and it narrows it in two separate ways.
The first is elimination. If you have learned 妈 and hear ma in a sentence about someone's family, the four tonal possibilities were never four. You needed to distinguish a real candidate from three non-words, which is a much easier judgement than choosing among four equally plausible sounds. Most of the accuracy you are enjoying comes from this and not from your ear.
The second is more interesting because it happens below the level of choice. William Ganong showed in 1980 that when a speech sound is ambiguous between two categories, listeners shift their perception towards whichever interpretation makes a real word. The bias operates on perception itself, not on the guess afterwards; the ambiguous sound is genuinely heard as the word-forming one. The same lexical pull has been reported for Mandarin tone, and it means a marginal contour on a familiar syllable does not feel marginal to you. It feels clear, and the clarity is supplied by your vocabulary.
This is why the experience is so hard to self-diagnose. You do not feel yourself guessing. You feel yourself hearing, and on the unfamiliar syllable you feel the hearing stop working, when in fact that is the first time it has been asked to do the job alone.

Why the score flatters you, and by how much
The size of the inflation depends on the material, and four common kinds of practice sit at very different points.
| Material | What the answer can be read off | What the score measures |
|---|---|---|
| A word from your current vocabulary list | The word, recalled from memory | Vocabulary, almost entirely |
| A known word inside a sentence | The word, plus the grammar around it | Comprehension, which is a real skill and not this one |
| A minimal-pair set on one syllable | Nothing outside the audio, but the options are visible | Discrimination between two named alternatives |
| A syllable you have never learned, alone | The audio and nothing else | Tone perception |
Nothing on that list is a bad exercise. They are measurements of different things, and the error is reading one of the top two as if it were the bottom one. A learner who is at 94 percent on their vocabulary deck and 71 percent on unfamiliar syllables does not have two contradictory scores. They have a vocabulary score and a tone score, and the tone score is the 71.
The gap also explains a specific frustration. Learners who have been at this a while report that their tones seem to have got worse as their Chinese improved. What actually happened is that their vocabulary grew, more of their listening became familiar, and the proportion of their practice that tested perception fell towards nothing. The skill did not decline. The measurement stopped taking place.
The honest test, and how to run it
A syllable that carries no meaning for you cannot be read off memory, so it puts the whole weight on the audio. Building that test takes about twenty minutes.
- Choose syllables you have never learned as words. Not words you find difficult, which are still words. Genuinely unknown ones, ideally with phonology you have no strong associations for.
- Present them alone. No sentence, no picture, no translation on the screen. Context is the thing being excluded.
- Answer once, quickly, and do not replay. Replaying gives your vocabulary time to search for a match, which is the exact process you are trying to shut out.
- Run at least eighty items and record which tone you pressed when you were wrong, not just how many you missed. The pattern of confusions is the useful output.
Expect the number to be lower and treat that as the point rather than as bad news. Whatever gap opens between this score and your vocabulary-deck score is a measurement of how much of your listening has been running on recall, and it is the most informative number in the whole exercise.
Once you have it, the confusion pattern tells you where to go next. If the errors cluster in one pair of tones, that is the ordinary plateau and it has a specific fix, described in why tone recognition stalls around 80 percent. If they are spread evenly across all four, the categories are still forming and more contrastive exposure is what you need.
How a drill corpus can enforce this
Tingyin is a Mandarin tone-listening trainer with exactly one exercise: a clip plays, you pick the tone. Every one of its 637 clips is a human recording carrying its source, speaker, licence and a hash of the audio, and the way those clips are distributed across the levels decides how much familiarity a learner can accumulate.
Level 1 is built for the opposite purpose and says so. It is 20 syllable families recorded in all four tones, 80 clips, so ba in the first tone and ba in the third differ in one dimension and nothing else. Repetition is the design. It is a place to build the categories, not a place to measure them, and the score you get there will be your highest.
Level 2 is the honest one, and the numbers make the difference concrete. Its 200 clips are drawn from 160 different syllables, and 127 of those appear exactly once in the whole level. There is not enough repetition for familiarity to build inside it. A syllable you meet once and never again cannot become a word you recognise, so the answer has to come out of the audio each time. Levels 1 and 2 are both free, and level 1 needs no account at all.
None of this involves any check on the pitch of a recording. Nothing in the project analyses the contour of a clip or certifies that a file carries the tone it is labelled with. The audio is trusted for its provenance, and the automated checks confirm the bytes, the licence and the hash. The corpus design above is about what a learner can memorise, not about verifying the tones themselves.

What to do with the gap
Two responses are available and one of them is wrong.
The wrong one is to stop using vocabulary knowledge. Using what you know to interpret what you hear is not cheating, it is what fluent listening consists of, and native speakers do it more aggressively than you do. In a noisy restaurant it is most of what keeps a conversation running, as the piece on hearing tones in noise goes into. The goal is not to disable the mechanism. It is to stop mistaking it for perception when you are trying to measure perception.
The right one is to separate the two exercises. Keep the vocabulary work for vocabulary, in whatever tool you already use for it, and keep a smaller block of unfamiliar-syllable listening for the perceptual side. They answer different questions and neither substitutes for the other.
There is a production version of this problem too, and it has the same cause. Learners produce accurate tones on rehearsed words and drift on new ones, because the rehearsed tone was stored with the word rather than assembled from the system. Hearing tones and saying them are separate skills, and both of them are flattered by familiar material in the same way.
Frequently asked questions
Why can I hear tones on words I know but not on new ones?
Because on a known word you are not identifying the tone from the audio, you are recognising the word and reading its tone out of memory. Knowing the word also eliminates three of the four possibilities before you decide, and lexical knowledge biases perception itself toward whichever interpretation makes a real word. On an unfamiliar syllable none of that is available, so the audio has to carry the whole judgement, which is the first honest measurement of your tone perception you have taken.
What is the best way to test my real tone accuracy?
Use syllables you have never learned as words, presented alone with no sentence and no translation, answered once without replaying. Eighty items is enough to see a pattern. Tingyin's second level is built this way, with 200 clips drawn from 160 different syllables and 127 of them appearing exactly once, so familiarity cannot accumulate inside the level. A vocabulary tool such as Anki or Pleco cannot give you this, because its entire purpose is to make the material familiar.
What is the difference between a tone problem and a vocabulary problem?
A tone problem persists on syllables that carry no meaning for you, and a vocabulary problem disappears the moment the word becomes familiar. Run the same set twice, once from your current word list and once from syllables you have never met, and compare the two scores. A large gap means your listening has been running on recall; a small gap with a low score in both means the perceptual categories are still forming.
How do I hear tones in Chinese names I have never met?
Slowly and worse than you hear anything else, which is normal rather than a personal failing. A name has no surrounding sentence to constrain it, often no meaning you can reach for, and frequently a syllable outside your vocabulary entirely, so every repair strategy you rely on elsewhere is unavailable at once. The practical move is to ask which character it is written with rather than asking for the sound again, because a repeat gives you the same information twice.
Is it safe to assume my drill score reflects my listening ability?
Only if the drill uses material you have not memorised. Any exercise built from your own vocabulary list measures recall and perception together and cannot separate them, and the proportion coming from recall grows quietly as your Chinese improves. That is why some learners feel their tones are getting worse over time when what actually happened is that their practice stopped testing the thing they wanted to measure.