Why Mandarin Tones Are Hard to Hear, Not to Say
Most English speakers learning Mandarin can produce a passable second tone within a week and still cannot tell a second from a third six months later. Production and perception are separate skills, and almost all study material trains the first one. What fixes the second is contrastive listening — hearing the same syllable in two tones, guessing, being told, and doing it a few hundred times.

The problem is in your ears, not your mouth
Ask someone six months into Mandarin what they find hard about tones and they will usually describe a production problem: my third tone comes out flat, my fourth is too soft. Watch the same person listen to a sentence at speed and the actual failure shows up somewhere else. They can say the tone. They cannot hear which one they were just told.
This is not a character flaw and it is not unusual. Adults learning a tonal language have to build a category that their first language never asked them to build, and building it takes exposure to contrasts, not explanations of contrasts. You do not learn to tell mā from má by reading that one is high and level and the other rises. You learn it by hearing them alternate until they stop sounding like the same syllable.
Tingyin is a Mandarin tone-listening trainer built around that one exercise: a clip plays, you pick which tone you heard, it tells you immediately. There are no characters to write, no grammar drills and no vocabulary lists — the whole application is the listening loop. Every level is free; a one-time purchase buys progress sync across devices and nothing else.
What English has already taught you to ignore
English uses pitch, heavily. It just uses it for something else. In English a rise across a word marks a question, a fall marks finality, and a jump marks emphasis — all of it applying to the phrase rather than the word. Your ear has spent decades learning to route pitch information to the part of your brain that handles attitude and sentence structure, and to discard it as a property of the word itself.
That habit is precisely what gets in the way. A rising second tone in the middle of a Mandarin sentence sounds, to an untrained English ear, like the speaker is asking something — so the pitch gets filed under intonation and the word identity is lost. The four tones and the neutral tone are not decoration on top of the syllable; they are part of its spelling.
| Tone | Shape | Example | Meaning |
|---|---|---|---|
| First | High and level | mā 妈 | mother |
| Second | Rising | má 麻 | hemp |
| Third | Low, dipping | mǎ 马 | horse |
| Fourth | Sharp fall | mà 骂 | to scold |
| Neutral | Short, unstressed, takes its pitch from what came before | ma 吗 | question particle |
Every chart like this one is true and almost none of it transfers. It is a description of what you are listening for, not practice at listening for it, and the gap between those two is where most learners spend a year.

The third tone is worse than the textbook says
Textbooks draw the third tone as a dip: down, then up again. In connected speech it very often is not. Before another syllable it usually stays low and never rises at all, and between two third tones the first one shifts to sound like a second — the sandhi rule every course mentions once and then leaves you to spot in the wild.
We have a measurement of how messy this gets, from an unexpected direction. Tingyin used to run an automated gate over every audio clip before shipping it: extract the pitch track, check the contour matches the tone the clip is labelled with, reject anything that fails. Against real recordings by native speakers, that gate rejected 93% of correct third tones. Not wrong recordings — correct ones, from people who have spoken Mandarin their whole lives, thrown out because a real third tone in a real word does not look like the shape in the diagram.
The gate has since been removed, because a test that fails 93% of the correct answers is measuring the wrong thing. But the number is worth carrying around as a learner, because it says something useful: if a pitch-tracking algorithm cannot match real third tones to the textbook contour, you are not going to do it by memorising the contour either. The only route in is hearing enough real ones that the category forms on its own.
Why it matters who is speaking
The same gate had a second result, in the other direction: it passed synthesised clips that sound wrong to anyone who speaks the language. That combination — rejecting good human recordings, waving through bad synthetic ones — is the argument against learning tones from text-to-speech, and it is worth knowing which one you are listening to when you pick a tool.
Tingyin ships 637 clips and every one of them is a recording of a person. 329 come from Yue Tan and 279 from Chen Wang through the audio-cmn corpus, with a further 29 from six named Lingua Libre volunteers, all under CC BY-SA 4.0 or CC0. Each clip carries its source, speaker, licence and a content hash, and the build checks that the bytes served are the bytes recorded. The vocabulary follows from the recordings rather than the other way round: a word only appears in a level if a verified human recording of it exists.
This is a constraint as much as a feature — it is why the corpus is 637 items across four levels rather than several thousand. For ear training that trade is the right way round. A thousand synthesised clips of a tone that is subtly wrong will train you to recognise something no Mandarin speaker produces.
What actually moves the needle
Four things, in rough order of how much difference they make:
- Guess before you are told. Passive listening does very little. The commitment — picking an answer, being right or wrong about it — is what turns exposure into learning. Every trial should end with you having taken a position.
- Work in minimal pairs. The same syllable in two tones, back to back, is worth far more than two unrelated words. It isolates the one variable you are trying to hear.
- Start with two tones, not five. Second against third is the pair that costs most learners the most time. Get that contrast reliable before adding the others back in.
- Keep it short and frequent. Ten minutes daily beats an hour on Sunday, because what you are building is a perceptual category and those form through repeated exposure rather than concentrated study.
None of that requires a particular app. A friend with a word list and patience will do it, and so will a deck of recorded pairs you assemble yourself. What is hard to arrange on your own is the volume — a few hundred trials a week, each with immediate feedback, on audio you can trust. That is the gap the training page exists to fill, and it costs nothing to find out whether it helps.
Frequently asked questions
What is the best way to learn to hear Mandarin tones?
Contrastive listening with immediate feedback: hear a syllable, commit to which tone it was, find out at once whether you were right. Tingyin is built around exactly that loop and nothing else, with 637 human recordings across four levels. Anki works too if you are willing to build the deck and source the audio yourself, and it gives you spaced repetition that a simple drill does not. What does not work is reading the tone chart again.
How long does it take to hear the difference between tones?
Weeks of daily practice for the easy contrasts, longer for second against third, which is the pair that defeats most people. The honest answer is that it depends on how many trials you get through rather than how many months pass. Ten focused minutes a day for a month is a few thousand judgements; an hour a week for the same month is a few hundred, and it shows.
Is it a problem that I can say the tones but not hear them?
It is the normal shape of the problem, not a sign you are doing something wrong. Production is a motor skill you can copy from a model, and perception is a category you have to build from exposure — the two come apart routinely, and most courses drill the first because it is easier to grade. It matters because comprehension is where the asymmetry bites: you can be understood while understanding very little.
What is the difference between a tone trainer and a full Mandarin course?
Scope. HelloChinese and Duolingo teach a whole language — vocabulary, characters, sentences — and tone practice is one component inside that. Tingyin does one exercise and nothing else, which makes it a supplement rather than a replacement: it will not teach you a single word of grammar. Pleco is worth having alongside either, as the dictionary you reach for when you meet a word in the wild.
Can I really not just learn the tone marks and work it out?
Not at conversational speed, no. The marks tell you which tone a syllable carries when you are reading it, which is a different task from identifying a tone in half a second of speech you did not expect. Tingyin exists because that second task needs its own practice, and the third-tone measurement above is the clearest evidence that the diagram and the sound have come apart: even a pitch-tracking algorithm could not reconcile them.