Mandarin Tones Survive a Phone Call
A phone call cuts the fundamental frequency of most voices, and it does not cut the tones. The channel passes 300 Hz to 3400 Hz, and what carries tone sits inside that. Tingyin drills 637 human recordings at full bandwidth, which is not what a call sounds like, and neither are Pleco's native-speaker recordings.

What the channel is, exactly
The telephone band is not a rumour or a rule of thumb. It is written down, and the documents are public. Four of them define the pipe a Mandarin syllable has to get through on a call.
| Standard | What it fixes | The number |
|---|---|---|
| ITU-T G.712 | The frequency range the transmission performance masks are specified over | 300 Hz to 3400 Hz |
| ITU-T G.711 | Sampling rate for digital telephone PCM | 8,000 samples per second |
| TS 26.071, AMR narrowband | Mobile narrowband input format and bit rates | 8,000 samples/s, 4.75 to 12.2 kbit/s, a new rate every 20 ms |
| ITU-T G.722.2, AMR-WB | Wideband mobile calls, sold as HD Voice | 16,000 samples/s, 6.6 to 23.85 kbit/s, 50 Hz high-pass |
Two things follow immediately. On a narrowband call the sampling rate is 8,000 per second, which is a hard ceiling regardless of what the microphone captured. And the floor is 300 Hz, which is above where a human voice actually vibrates.
The fundamental does not make it through
In the Mandarin speech Fei Peng and colleagues used for a 2018 study of how the brainstem represents tone contours, the fundamental frequency across all their material ran from 80 Hz to 180 Hz. Every one of those values is below the 300 Hz floor of the telephone mask. The frequency that is the pitch you are listening for is removed before the call leaves the handset.
This should be fatal and is not, for a reason that has already been covered here in the context of small speakers: your auditory system reconstructs the missing fundamental from the harmonics above it. The interesting question is not whether that works, since it demonstrably does. It is which harmonics have to survive for it to work, and whether the telephone band happens to keep them.

Why the tones get through anyway
The harmonics of a voice split into two groups by whether the ear can separate them. The low-numbered ones each get their own auditory filter and are called resolved; above roughly the tenth harmonic they start sharing filters and are called unresolved. Peng and colleagues needed to isolate the resolved set for their experiment, and for speech with a fundamental between 80 and 180 Hz they did it with a low-pass filter at 800 Hz.
That is the whole answer. The resolved harmonics of a normal speaking voice live below about 800 Hz, and the telephone band starts at 300 Hz and runs to 3400 Hz. Most of that region arrives intact. Their summary of the literature on what listeners do with it is that tone identification accuracy for sounds containing the low resolved harmonics is nearly perfect.
The contrast makes the point sharper. Kong and Zeng found that tone identification from synthetic stimuli containing the fundamental and its harmonics was nearly perfect even with white noise at 0 dB signal-to-noise ratio, while the same listeners scored around 60% on stimuli carrying only the amplitude envelope of the speech at the same noise level. The phone keeps the first kind of information. Whatever is going wrong on a bad call, it is usually not the tones.
What the ceiling actually takes
The damage is at the other end of the band. Consider what it takes to tell a dental sibilant from a retroflex one, the s from the sh, the z from the zh, the c from the ch. When Yu-Ying Chuang and colleagues measured that distinction across 11,364 sibilant tokens from 331 speakers, the first thing they did was filter out everything below 1000 Hz before computing the centroid frequency, because nothing below that line helps. The information that separates those consonants sits in the upper part of the spectrum, which is the part a 3400 Hz ceiling trims and a 300 Hz floor does not protect.
Now put a name in that channel. A name is the one thing on a call with no redundancy: no surrounding sentence to constrain it, no topic to guess from, often no word you have heard before. Everything else you can repair from context. When someone says 张 or 常 or 陈 down a phone line, the tone probably arrived and the consonant probably did not, and there is nothing else in the utterance to reconstruct it from. That is why asking a stranger to repeat a surname three times feels like a tone problem and generally is not one.
If the call is on AMR-WB, sold under names like HD Voice, the arithmetic changes. 16,000 samples per second instead of 8,000, and a high-pass at 50 Hz rather than 300 Hz, which puts the actual fundamental of most voices back inside the channel and lifts the ceiling well above where the sibilant information lives.
What to do about it
Tingyin is a Mandarin tone-listening trainer: a clip plays and you pick the tone. The 637 clips are human recordings, each carrying its source, speaker, licence and sha256, and they are served at full bandwidth rather than through anything resembling a phone line. Levels 1 and 2 are free, and level 1 needs no account. That is a deliberately clean signal, and the gap between it and a call is the subject of this article rather than a defect in either.
- Stop treating a bad call as a tone failure. If you caught the melody and lost the word, the loss was segmental. Ask for the character, not for the tone.
- Ask for a name by its parts. The written form is the repair channel the audio does not have: the 弓 in 张, or which 陈. Native speakers do this with each other constantly and it is not a learner's crutch.
- Prefer a data call for a hard conversation. Voice over an app is frequently wideband where a circuit-switched call is not, and the difference is the consonants rather than the pitch.
- Keep drilling tones on clean audio. Deliberately degrading your practice material to match a phone would remove the sibilant cues, which the drill was never testing, and leave the tonal cues, which it was. Hearing the four tones apart is a separate skill from parsing a bad line, and only one of the two improves by practising it.

Frequently asked questions
What is the best way to practise Mandarin for phone calls?
Practise the tones on clean audio and practise the repair strategies separately. The tonal information survives the 300 Hz to 3400 Hz telephone band because it rides on harmonics that fall inside it; what you lose is consonant detail, and no amount of listening drill replaces knowing how to ask which character someone means. Tingyin covers the first half. Pleco is useful for the second, since you can look up a name by component while the person is still on the line.
How do I catch a Chinese name on a phone call?
Get the tone first, because that part probably arrived, then ask for the character rather than a repeat of the sound. Repeating an unfamiliar surname down the same degraded channel gives you the same degraded information a second time. Asking which component it is written with changes the medium, which is the only thing that adds information.
Is it safe to assume HD Voice fixes the problem?
It fixes a good part of it when both ends and the whole path support it, which you cannot verify from your side of the call. AMR-WB samples at 16,000 per second against 8,000 for narrowband, and high-passes at 50 Hz rather than sitting behind a 300 Hz floor. A call that starts wideband can also drop to narrowband mid-conversation without telling you.
What is the difference between a phone speaker and a phone line?
They fail at opposite ends of the spectrum. A phone loudspeaker cannot reproduce the low fundamental but passes the harmonics fine, which is a playback problem covered in the piece on headphones for tone practice. A phone line removes the fundamental by specification and also cuts everything above 3400 Hz, which takes consonant detail the speaker would have reproduced happily.
So am I actually mishearing tones on calls, or not?
Mostly not, and the feeling that you are is worth interrogating. The channel keeps the resolved harmonics that tone identification runs on, and the literature has listeners near ceiling on those even in noise. What a call takes is the fine consonant detail and the visual context, and the sensation of both going at once is very easy to file under "my tones are bad" when the tones were the part that got through.