Chinese Tone Practice: How to Drill Your Ear, and What Doesn't Work
Tone practice that works is forced-choice identification with immediate feedback: hear a syllable, commit to an answer, be told at once whether you were right. Charts, unchecked repetition and vocabulary lists do very little for perception. Tingyin drills the identification half only - it cannot hear you speak, so pair it with a tutor or italki for production.

What most people mean by tone practice
Ask ten learners what they do for tone practice and you get roughly three answers. They look at the contour chart - the one with the four arrows. They repeat after a recording. They learn vocabulary with the tone marks attached and hope the tones come along for the ride.
People genuinely do all three, some of them daily, for months. And a great many of those people are still guessing between the second and third tone a year later. That is not a motivation problem. It is that none of the three is practice at the thing they actually want to be able to do, which is hear an unfamiliar syllable go past at speed and know which tone it carried.
The useful frame is that tone is two separate skills sharing one name. Producing a tone is a motor task: there is a model, you copy it, and you can feel roughly whether you matched. Identifying a tone is a perceptual category - the same sort of thing that lets you hear the difference between the vowels in bit and bet without thinking about it. Categories like that are not built by explanation. They are built by exposure to contrast, repeatedly, with the answer supplied. That split is the subject of why Mandarin tones are hard to hear, and it is the whole reason a listening drill is a different kind of product from a course.
Tingyin is that drill and nothing else: a clip plays, you pick which tone you heard, it tells you immediately. No characters, no grammar, no vocabulary list. Level 1 is free with no account at all, level 2 is free once you sign in, and levels 3 and 4 are paid.
The four things that actually build the category
In rough order of how much difference each one makes:
- Commit to an answer before you are told. A forced choice is what separates practice from listening. Guessing feels uncomfortable and that discomfort is the mechanism - being wrong about a specific prediction is what sharpens the boundary. Passive exposure with the answer already visible does almost nothing.
- Get the answer within a second. Feedback delayed by even a few seconds is much weaker, because the sound you were judging has already faded. This is the single reason a drill beats a podcast for this particular job.
- Work in minimal pairs. The same syllable in two different tones, back to back, isolates the one variable you are trying to hear. Two unrelated words in two different tones give your ear a dozen other cues to latch onto instead - vowel, length, consonant - and it will happily use them.
- Keep sessions short and frequent. Ten minutes a day beats seventy minutes on Sunday, and it is not close. A perceptual category consolidates between sessions rather than during them, so the number of sessions matters more than the length of any one of them.
Level 1 in Tingyin is built exactly on the third of those: 80 single-syllable clips arranged as minimal pairs, the same syllable presented in each of its tones. It is deliberately the narrowest exercise in the app, because the narrow version is where the contrast is cleanest.
One thing worth adding that no fixed corpus can give you: variety of speakers. Tones sit differently in a low voice and a high one, in fast speech and careful speech, and a category built on a handful of voices is more brittle than one built on many. Keep some unstructured listening in your week - a podcast, a show, a conversation partner - alongside whatever drill you use.
What does not work, and why it feels like it should
The failure mode most learners fall into is not laziness. It is doing something adjacent to the target skill and assuming the transfer will happen.
| What you do | What it trains | Builds identification? |
|---|---|---|
| Study the contour chart | Knowing what the tones are | No |
| Repeat after a recording, alone | Production, with nobody checking | Barely |
| Learn vocabulary with tone marks | Reading, and recall of the written tone | No |
| Shadow with a teacher correcting you | Production, checked | Some |
| Forced-choice identification with feedback | Perception | Yes |
The chart is the most seductive of these because it is genuinely true and it takes ninety seconds to read. But knowing that the third tone dips is a fact about the tone, not an ability to spot one. Worse, the chart describes an isolated, stressed syllable, which is close to the rarest environment a third tone actually occurs in - the rest of the time it either stays low and flat or turns into something that sounds like a second tone, as the sandhi rules describe.
Repeating after a recording with nobody listening is the second trap. It feels like the most serious kind of work - you are producing Mandarin out loud - but the loop has no error signal in it. You compare your output to the model using the same ear that cannot yet tell the tones apart, so a systematic mistake gets rehearsed rather than corrected. That is not an argument against speaking practice. It is an argument for having someone who can hear you.
And learning words in isolation with the tone marks attached trains a lookup table, not an ear. You end up able to answer what tone is měi? instantly while still failing to catch it in a sentence, because the two tasks share a name and nothing else.

What ten minutes should actually look like
A session that does the job has a shape. Roughly:
- Two minutes on a pair you find hard. For most English speakers that is second against third. Only two options on screen, nothing else, until the hit rate stops being embarrassing.
- Five minutes on all four, single syllables. Mixed order, no warning about which is coming. This is the bulk of the work.
- Three minutes on two-syllable words. A tone inside a word behaves differently from a tone alone, and the transfer between them is smaller than you would expect.
That last step is the one people skip, and it is where most of the real difficulty lives. Tingyin's corpus is 637 clips, and the distribution says something about where the problem is: 280 clips are a single syllable, 237 are two, 107 are three and 13 are four. Sandhi is applied in 55 of them and 80 contain a neutral tone - and every one of those is at least two syllables long, because a neutral tone cannot exist on its own.
There is a detail in how those clips are labelled that matters more than it sounds. The tone recorded against each syllable is the tone as spoken, not as written. A word written 3+3 is stored as 2+3, because that is what a speaker says and therefore what you have to identify. Across all 637 clips there is not a single one labelled third-then-third. If a drill marks you wrong for hearing what was actually said, it is training you to answer a spelling question instead of a listening one.
The other thing worth knowing about a corpus you are going to spend hours inside is where the audio came from. Every one of Tingyin's 637 clips is a recording of a person rather than a synthesised one, and each carries its source, speaker, licence and a content hash that the build checks against the bytes actually served. It is not glamorous, but a drill is only as good as its answer key, and a clip that does not sound like what it claims to be actively teaches you the wrong boundary.
The half a drill cannot reach
Be clear about the limit. Identification training does not fix your production. Tingyin never hears you - there is no microphone in the exercise, no pronunciation score, no feedback on anything you say. It can tell you that you failed to recognise a fourth tone. It cannot tell you that yours comes out too gently.
For that you need a person, and there is no honest substitute. An italki tutor at a low hourly rate, a language exchange partner, a teacher in a class - someone whose ear already has the categories and who will interrupt you. Automated pronunciation scoring exists in several of the big apps and is worth roughly what you would expect from a machine grading a skill it cannot fully measure: useful as a rough flag, misleading as a verdict.
The two halves do help each other, and in a specific direction. Perception generally leads: once you can reliably hear a contrast, correcting your own production becomes a much smaller problem, because you can finally tell when you have got it wrong. That is the argument for putting a few weeks of pure listening work in early, before you have rehearsed a wrong second tone four thousand times. If you want a sense of how the timelines actually run, how long it takes to learn Chinese by skill breaks it down; if you are earlier than that and cannot yet segment what you hear, listening practice below the sentence is the place to start.
You can find out whether any of this helps for nothing. Level 1 needs no account at all - 80 single-syllable minimal pairs, which is more than enough to tell you whether forced-choice drilling suits you. Level 2 opens once you sign in. Levels 3 and 4 are paid, at $5.99 a month or $15.99 once on the web, and paying also adds progress sync across devices and offline downloads. The web is the only place to pay today: the native apps are not published yet.
Frequently asked questions
What is the best way to practise Chinese tones?
Forced-choice identification with immediate feedback, in short daily sessions, starting with minimal pairs of the same syllable. Tingyin does only this - 637 clips across four levels, level 1 free without an account. Anki can do it too if you are prepared to build the deck and source audio yourself, and it adds spaced repetition. What reliably does not work is re-reading the tone chart, however many times you re-read it.
Can Tingyin tell me whether my pronunciation is correct?
No. There is no microphone anywhere in the exercise. Tingyin plays a clip and you pick the tone you heard, so it trains recognition and is silent about production. For pronunciation you need someone who can hear you - an italki tutor, a language partner, a teacher. HelloChinese includes automated speech scoring if you want a rough machine check, but treat it as a flag rather than a verdict.
How much tone practice per day is enough?
Ten focused minutes daily does more than an hour once a week, because perceptual categories consolidate between sessions rather than during them. What matters is the number of judgements you make, not the minutes on a timer: ten minutes of forced-choice drilling is a few hundred trials with feedback, where ten minutes of passive listening is zero. Consistency beats intensity here more than in almost any other part of learning Mandarin.
Should I practise tones in isolation or in words?
Both, in that order. Isolated syllables give the cleanest contrast and are where the category forms; words are where it has to survive, because tones shift in connected speech and neutral tones only exist attached to something. Tingyin's levels move through exactly that ladder - 280 single-syllable clips, then 237 of two syllables, 107 of three and 13 of four. Do not stop at the isolated ones.
Is a tone trainer a replacement for a Mandarin course?
No, and it is not trying to be. Tingyin teaches no characters, no grammar and no vocabulary - it is one exercise. Duolingo and HelloChinese teach a whole beginner course and include some tone work inside it; Pleco is the dictionary you will want alongside either. A tone trainer is a supplement for one specific weakness, which is worth the ten minutes precisely because a full course cannot give that weakness enough attention.