Hearing Mandarin Tones and Saying Them Are Different Skills
Listening practice does improve how you say Mandarin tones. In one study, American learners who trained only their ears had their productions judged 18% more accurately by native listeners, with no production training at all. Tingyin trains only the ear, one clip at a time. HelloChinese does speech recognition if you want the other half.

The transfer is real and it has been measured
The reasonable objection to an ear-training app is that hearing a tone and producing one are different motor problems, so drilling the first should do nothing for the second. Yue Wang, Allard Jongman and Joan Sereno tested exactly that and published the result in the Journal of the Acoustical Society of America in 2003.
Sixteen American learners of Mandarin were recorded reading a word list. Eight of them then went through eight sessions of perceptual tone training over two weeks; the other eight did nothing. Everyone was recorded again. Those recordings went to Beijing, where 80 native Mandarin judges, screened first on a native-produced list, wrote down which tone they heard. Nobody in the study ever received feedback on their own speech.
| Group and material | Before | After | Change |
|---|---|---|---|
| Trained, words used in training | 57% | 75% | +18 |
| Trained, words never used in training | 55% | 68% | +13 |
| Control, words used in training | 57% | 61% | +4 |
| Control, words never used in training | 58% | 59% | +1 |
The authors state the finding without qualification: the learners' productions of Mandarin tone improved without any production training. It also generalised, showing up on words the trainees had never heard during training, which rules out the explanation that they simply memorised how eighty specific words were supposed to go.
The spread underneath the average is worth knowing before you set expectations. Across the eight trainees, starting accuracy ran from 24% to 77%, and improvement ran from 3% to 31%.
Where the transfer stops
Now the part that an ear-training product has an incentive not to mention. The improvement was not spread evenly across the four tones, and the tone it failed on is the one learners most want fixed.
In the perception half of the same research programme, the third tone was relatively good to start with and improved significantly with training. In production it was poor at the pretest and stayed poor at the post-test. The paper's own summary of the pair of results: while tone 3 was relatively easy to identify, it was difficult to produce, and was resistant to improvement. Across both tests, the trainees' third tone was significantly worse than their first, second and fourth.
The acoustic analysis found the same split along a second axis. Post-training contours moved measurably closer to the native norm, and pitch height moved less than pitch contour did. Shape is learnable by ear. Register, the absolute height a speaker puts a tone at, is stickier. If you want the fuller account of why the third tone in particular resists this, the third tone is a different shape than it is taught as covers it.

The mirror error
The myth runs in both directions, and the second version is the more common one among people who have been studying a while: I can produce all four, so I must be able to hear all four.
At the pretest, the two error patterns were almost the same shape. The percent errors for perception and production across the six tone pairs correlated at r = 0.98, and the rank order of which pairs were hardest correlated at 0.94. Second and third were the most confused pair in both. So far this supports the intuition.
The direction of the confusion did not match. In perception, a second tone was misheard as a third more often than a third was misheard as a second. In production, a third tone was produced as a second more often than the reverse. The same pair, failing opposite ways depending on which end of the mouth you are standing at. Knowing you have a second-and-third problem does not tell you which of the two skills is broken, and fixing one does not automatically report on the other.
What an ear-training app does and does not do
Tingyin is a Mandarin tone-listening trainer, and listening is the whole of it: a clip plays and you pick the tone. There is no recording, no pronunciation score and no speech recognition anywhere in it. Every clip is a human recording carrying its source, speaker, licence and sha256, 637 of them. Levels 1 and 2 are free, and level 1 works without an account.
Set against the study, that scope buys a specific and defensible thing. Perceptual training is the intervention that produced an 18% production gain in people who never practised producing, and it is a task you can do on a phone, in silence, without a partner and without embarrassment. It is not a substitute for the parts it does not touch, and the paper is clear about which those are: the third tone, and pitch height generally.
The order that follows from the evidence is not controversial. Ear first, because perception appears to lead production rather than the other way around; production practice with feedback after, aimed at the specific things the ear does not carry across. Hearing the difference between the four tones is the prerequisite, and it stops being the bottleneck sooner than most learners expect.

Frequently asked questions
What is the best way to fix my Mandarin tones, listening or speaking?
Listening first, and the evidence for that ordering is direct: eight sessions of perceptual training over two weeks raised native listeners' identification of the trainees' own speech by 18%. Then production practice with feedback, because the same study found the third tone and pitch height did not come along for the ride. Tingyin covers the first half only; HelloChinese includes speech recognition and covers reading, writing, speaking, vocabulary and grammar alongside it.
How do I know whether my problem is hearing or saying?
Test them separately, because the error patterns look alike and fail in opposite directions. A tone identification drill tells you about perception. Recording yourself and having a native speaker label what you said, with no context and no lookahead, tells you about production. In the 2003 study the same tone pair was the hardest in both, but second tone drifted toward third when listening and third drifted toward second when speaking.
Is it safe to rely on an app that never listens to my voice?
It is a smaller privacy surface, since an app with no microphone permission cannot record you, and Tingyin does not ask for one. The trade is coverage rather than risk: nothing in an ear-training drill can tell you your third tone is landing as a second, which is the most common production error the study found and the one perceptual training did not reach.
What is the difference between tone perception and tone production training?
Perception training gives you a clip and asks which tone it was, with immediate feedback. Production training records you and evaluates what came out, either by software or by a teacher. They are correlated but not interchangeable: perception training reliably moves production, at 18% in the measured case, while there is no symmetric result showing that drilling production fixes hearing.
Will listening practice alone actually make me sound better?
Measurably, yes, though less than a full course would and unevenly across the four tones. The honest range is the one from the study: eight learners, improvements from 3% to 31%, with the biggest gains for the people who started lowest. If your tones are already good enough to be understood and you are chasing the last stretch of accent, this is the point where a teacher earns their fee.