Hearing Mandarin Tones in Noise: What Goes First
Noise does not erase Mandarin tones evenly, and plain hiss barely touches them at all. What wrecks tone perception is other voices, because a competing talker brings a pitch contour of its own into exactly the frequency band your ear reads tone from. The third tone goes first, for two reasons that compound: it spends more of itself well below its own peak than any other tone, and the feature that identifies it sits inside that quiet stretch.

The narrow band tone actually lives in
A Mandarin tone is a pitch contour, and the pitch of a speaking voice comes from the rate the vocal folds vibrate. In the Mandarin speech Fei Peng and colleagues used in 2018, that rate ran from 80 Hz to 180 Hz across all their material. But the ear does not read pitch off the fundamental alone. It reads it off the low harmonics, the ones each landing in an auditory filter of their own, and for voices in that range those resolved harmonics sit below roughly 800 Hz.
So the whole of the tonal signal is packed into a band a couple of octaves wide, at the bottom of the speech spectrum. That is a small target, and whether a given noise damages your tone perception depends almost entirely on how much energy that noise puts inside it. This is a different question from the one the article on phone calls answers. A phone removes a region of the spectrum by specification and passes the rest cleanly. Noise leaves the whole spectrum in place and buries part of it. Removal and masking are separate failures and they cost you different things.
Why hiss is nearly harmless
The result that surprises people is how much broadband noise tone perception survives. Kong and Zeng found that identification of tones from synthetic stimuli carrying the fundamental and its harmonics stayed near perfect with white noise at a 0 dB signal-to-noise ratio, which is noise as loud as the speech. Strip those stimuli down to the amplitude envelope alone and the same listeners fell to around 60 percent at the same noise level.
The reason is that periodicity is a redundant cue. A voice at 120 Hz puts energy at 120, 240, 360, 480, 600 and so on, all rising and falling together in lockstep, and the auditory system only needs a few of those to survive to recover the pattern. Flat noise spreads itself thin across the entire spectrum, so it never fully covers any one of them. Extractors that work from redundant, correlated evidence degrade slowly, and this one degrades very slowly indeed.
That explains a fan, an air conditioner, road roar, rain. They are loud, they are annoying, and they take far less of your tone perception than they seem to. The trouble starts when the interference stops being noise in the technical sense and becomes another voice.
What actually breaks it
A competing talker is a bad masker for tone because it is not a masker at all in the simple sense. It is a second harmonic stack with its own fundamental, its own contour, and its own onsets and offsets, moving in the same low band and behaving in every respect like a signal your ear is built to track. There is nothing for the periodicity extractor to reject.
| Interference | Where its energy sits | Damage to tone |
|---|---|---|
| Fan, air conditioning | Broad, weighted low, but flat and steady | Small, even when it is loud |
| Traffic, road roar | Concentrated low, mostly below the voice | Small to moderate, worse for a low-voiced speaker |
| One person talking nearby | A rival harmonic stack in the same band | Large, and largest when their voice is near your talker's pitch |
| Restaurant babble, several tables | Many rival stacks, plus a broadband floor | Large, and the hardest to attend around |
| Music with a sung vocal | A pitch contour designed to be followed | Large, out of proportion to its volume |
Two practical consequences follow. First, the loudness of the interference is a poor predictor of how much it hurts, which is why a quiet cafe with two nearby conversations can be worse than a loud one with a wall of undifferentiated sound. Second, a talker whose voice sits far from your target's in pitch is easier to ignore than one that sits close, so a room of speakers similar to the person you are listening to is close to the worst case.

Why the third tone goes first
Tingyin is a Mandarin tone-listening trainer: a clip plays, you pick the tone. Its 637 clips are human recordings that each carry their source, speaker, licence and a hash of the audio, and the manifest binding them records the duration and the loudness of every file. That gives two hard numbers relevant here, and neither is a measurement of pitch, which nothing in this corpus checks.
Across the 280 single-syllable clips, 70 per tone, the third tone averages -19.90 dBFS against -15.93 for a first tone. Roughly four decibels. All four tones peak within 0.13 dB of each other, so that gap is not a difference in how loud a third tone gets. It is a difference in how much of the syllable is spent well under its own peak, and a third tone spends more of itself down there than any other.
The second number says where. A third tone is the longest of the four, averaging 1.330 seconds against 1.012 for a fourth tone in the same corpus, and it is long precisely because it dips and comes back. So the extra time it spends below peak is the dip, and the dip is the part that identifies it. The bottom of a pitch contour is where a voice is quietest, often creaky, and sometimes not fully voiced at all. The single feature separating a third tone from a second is located in the weakest moment of the syllable, and in a noisy room it is the first thing to go under.
Put those together and the failure mode is predictable. In noise the third tone stops being heard as a third tone and starts being heard as a second, because the dip disappears into the floor and all that survives is the rise on the way out. This is the same pair that dominates quiet-room errors as well, which is covered in why tone recognition stalls around 80 percent. Noise does not create a new weakness. It widens the one you already had.
The full-dip third tone is also the citation form, the one spoken alone and deliberately. In connected speech it is usually shortened to a low tone that never comes back up, which the article on the third tone goes through properly. A low, short, quiet syllable in a loud room is close to the least recoverable thing Mandarin produces.
The order things go
As a room gets worse, the contrasts do not fail together. They fail in a sequence, and knowing the sequence tells you what your own mishearing means.
- The neutral tone goes first, and it barely counts as a casualty. It is short, quiet and unstressed by definition, and it disappears at noise levels that leave everything else intact.
- Third against second goes next, for the two reasons above. The turning point is lost before the overall direction is.
- First against second goes after that. Both stay high, and separating them needs the slope of a contour that is now only partly visible.
- The fourth tone is the last one standing. It is a large fall from the top of the range, over quickly, and a big fast movement is exactly the thing that survives partial evidence.
Note what is not on that list: loudness. It is tempting to expect the quietest tone to be hardest and the strongest to be easiest, but the four tones peak within 0.13 dB of each other in this corpus, and loudness carries no tone information even in a silent room. What the third tone's low average level costs you in noise is signal-to-noise ratio, not a missing cue.

What to do about it
The first thing is to stop training in noise on purpose. Adding babble to your drill audio makes the exercise harder without making it more informative, because the trials it costs you are concentrated in one pair you probably already know is weak. Difficulty and usefulness are not the same variable, and the version of this drill that improves fastest is the clean one where the feedback is unambiguous. Tingyin serves its clips at full bandwidth with nothing added, and levels 1 and 2 are free, with level 1 needing no account at all.
The second is to change what you ask of a noisy room. In a restaurant, the recoverable information is the syllable and the context, not the contour of a low syllable at the end of a phrase. Native listeners are not extracting those either; they are reconstructing them from the word, which is the same repair strategy that makes an unfamiliar name so much harder than a familiar one, discussed in tone accuracy on words you do not know.
The third is positional and it is worth more than it sounds. Sit so the person you want to hear is not in line with the loudest competing talker, because your ability to separate two voices depends heavily on them arriving from different directions. That advantage costs nothing, survives a loud room, and disappears entirely on a speakerphone, where every voice arrives from the same point.
Frequently asked questions
What is the best way to practise hearing Mandarin tones in noise?
Practise on clean audio and treat the noisy room as a separate skill. Tone identification holds up remarkably well against steady broadband noise, so adding hiss to your drill adds difficulty without adding information, and what actually defeats you in a restaurant is competing voices rather than volume. Tingyin serves its clips at full bandwidth for that reason. If you want practice against real competing speech, a podcast or Du Chinese audio played in a genuinely busy place is a better simulation than any noise-added drill.
Why is the third tone so hard to hear in a restaurant?
Two things stack against it. All four tones peak at the same level in Tingyin's 280 single-syllable clips, but the third tone averages about four decibels lower across the whole syllable, which means it spends more of itself well under its own peak than any other tone. That extra quiet stretch is the dip, and the dip is exactly what separates a third tone from a second. Lose it to the noise and what remains is a rise, which is a second tone.
How do I know whether noise or my ear is the problem?
Run the same material twice, once in a quiet room with headphones and once in the place that is giving you trouble, and compare which tones you got wrong rather than how many. If the errors are spread across all four tones in both conditions, the room is not the variable. If they concentrate on the third tone only when it is noisy, that is the masking pattern and it is normal.
What is the difference between noise masking and a phone line?
A phone line removes part of the spectrum permanently and passes the rest cleanly, while noise leaves everything in place and buries some of it. The telephone band cuts below 300 Hz and above 3400 Hz, which takes consonant detail and leaves the harmonics tone rides on intact. Noise does the opposite when it comes from other voices, because a competing talker puts energy exactly where the tonal information is. Different failure, different repair.
Is it safe to assume noise-cancelling headphones will help?
For steady low-frequency noise like a plane cabin or an air conditioner, yes, and that is what the technology is good at. For a room full of talking it helps far less, because active cancellation works on predictable periodic sound and speech is neither, and the interference that damages tone perception most is precisely the speech. Expect a large improvement on a train and a modest one in a cafe.