Why Mandarin Tones Are Harder on a Phone Speaker
A Mandarin tone is a movement in the fundamental frequency of a voice, and for most adult speakers that fundamental sits between roughly 85 and 255 Hz. A phone loudspeaker is a few millimetres across and reproduces almost nothing down there. You still hear the pitch, because your ear reconstructs it from the harmonics above, and most of the time that works well enough. Where it stops working is the bottom of the third tone and any room with noise in it, which are exactly the two places learners conclude they are tone deaf.

The thing you are listening for is lower than your phone can play
Strip a syllable back and the tone is one measurement over time: where the vocal folds are vibrating, and which way that number is moving. First tone holds it level and high. Second tone raises it. Fourth drops it fast. Third takes it down and, in isolation, back up.
That number is the fundamental frequency, and it is low. A typical adult male voice runs somewhere around 85 to 180 Hz, a typical adult female voice around 165 to 255 Hz. Now consider what is playing it. The speaker in a phone is a driver a few millimetres across in a sealed slot with no cabinet behind it, and physics does not let something that small move enough air to make a 120 Hz tone at any useful level. Below a few hundred hertz its output falls off a cliff.
So when you play a drill clip through a phone speaker, the frequency that is the tone is not really coming out of it.
Why you can usually hear it anyway
Because a voice is not one frequency. Vocal folds vibrating at 120 Hz also produce energy at 240, 360, 480 and on up, and those harmonics are well inside what a small speaker can manage. Your auditory system takes that ladder of harmonics and works backwards to the spacing that would produce it, which is the fundamental, and hands you a pitch that is not physically present in the sound.
This is the missing-fundamental effect, and it is the reason a pocket radio can play a bass line at all. It also means the tone contour survives the trip: if the fundamental rises, all the harmonics rise with it, so the shape is still there in what reaches you. Most of the time you hear the tone correctly on a phone speaker and there is nothing to fix.

Where the shortcut runs out
Two places, and they are not evenly distributed across the four tones.
The bottom of the third tone. A third tone in running speech mostly does not rise back up; it goes down to the bottom of the speaker's range and stays there, often slipping into creak, where the folds are vibrating slowly and irregularly. That is the weakest, lowest-energy part of the whole syllable, its harmonics are faint, and it is the part a small speaker gives you least of. What the third tone actually does covers the shape itself; the point here is that it is the tone your equipment degrades first.
Any room with noise in it. The reconstruction depends on hearing enough of the harmonic ladder. Put a bus, a café or a fan on top of it and the quieter harmonics get masked, the ladder gets patchy, and the pitch you were handed for free becomes work. Practising without a connection notes the same thing from the other end: background noise makes the third tone harder, and a train is a bad place to find that out.
| Listening setup | Fundamental reaching you | Where it starts to hurt |
|---|---|---|
| Phone speaker, quiet room | Almost none, reconstructed from harmonics | Low third tones, and only sometimes |
| Phone speaker, noisy room | Almost none, with the harmonics masked too | Everything below your comfortable range |
| Laptop speakers | A little more, still rolled off | Third tone, quiet clips |
| Earbuds or headphones | The actual fundamental, in your ear canal | Nothing equipment can be blamed for |
What to change, in order
- Use headphones or earbuds for drilling, even cheap ones. A driver sitting in your ear canal does not have to move much air to deliver 120 Hz, which is the whole reason it can.
- Turn it up further than feels necessary. The low end is the first thing to disappear at low volume, and it is the part carrying the tone.
- Drill somewhere quiet, and keep the noisy places for listening you already find easy. Noise takes the harmonics, and the harmonics are the backup.
- If a specific clip keeps beating you, replay it on headphones before deciding anything about your ear. Nine times in ten the second listen is not close.
What is actually in the clips
Tingyin is a Mandarin tone-listening trainer, and the whole exercise is one thing: a clip plays, you pick the tone. The material behind that is 637 recordings, every one of them a human recording carrying its source, speaker, licence and sha256 in a manifest that ships with the app. Eight speakers appear in it, though two of them do most of the work: Yue Tan on 329 clips and Chen Wang on 279, which is 608 of the 637 between them. Licensing is 622 under CC BY-SA 4.0 and 15 under CC0.
The clips are short by design, 0.89 to 2.10 seconds with a median of 1.20, and 280 of them are single syllables. Those 280 are split exactly evenly by tone, 70 each, so a run of first-tone answers is chance rather than a loaded deck. Level 1 is free and needs no account, and level 2 is free once you are signed in.
None of that helps if the sound arrives through a slot in a phone case. The clips are the part we can control; the last few centimetres are yours.
Frequently asked questions
What is the best way to listen to Mandarin tone drills?
Headphones or earbuds, in a quiet room, louder than you think you need. Tingyin's clips are human recordings between roughly one and two seconds long, so there is not much material to work with and losing the low end of it costs you more than it would in a long sentence. Pleco's dictionary audio has the same property for the same reason, and it is worth the same treatment.
How do I tell whether it is my ear or my speaker?
Take a clip you got wrong and play it again on headphones without changing anything else. If it is suddenly obvious, that was equipment. If it is still ambiguous, that is a genuine perception gap and drilling is the answer. Do this before you draw any conclusion about yourself, because the tone-deaf question almost never has the answer people assume.
Is it safe to practise tones with a Bluetooth speaker?
It is better than a phone speaker and worse than headphones, and it depends entirely on the speaker. A small portable one has the same physical limit as a phone and adds a room to the problem. Nothing is being damaged either way; you are just working harder for less, which is the argument against it.
What is the difference between hearing a tone and identifying it?
Hearing it is the signal getting to you intact, which is what equipment decides. Identifying it is mapping the contour you heard onto one of four categories, which is what practice builds. Confusing the two is expensive: a learner who is losing the signal spends months drilling a skill that was never the problem. Apps that teach tones as part of a general course, like HelloChinese and Duolingo, rarely separate the two either.
So my phone has been sabotaging me this whole time?
Mostly no, which is the honest answer. The harmonic reconstruction is good and you have been hearing the tones correctly most of the time. It is a narrow effect that shows up on the low end of the third tone and in noise, and those two cases are common enough in a real drill session to matter. Swapping to earbuds costs nothing and removes the question.