Mandarin Tones vs English Stress: Loudness Is Not a Cue
Loudness carries no tone information. Across the 280 single-syllable recordings Tingyin plays at levels one and two, all four tones peak within 0.13 dB of each other. Average level does differ, and it runs opposite to the folk rule: the fourth tone, the one learners are told is the strong one, is the quietest of the three short tones even when you compare clips of the same length. What a fourth tone has is a fall in pitch and an early peak in energy, not more volume.

Where these numbers come from
Tingyin is a Mandarin tone-listening trainer: a clip plays, you pick the tone, and that is the whole exercise. Every clip it ships is a human recording carrying its source, its speaker, its licence and a hash of the audio, and the manifest that binds them also records two loudness figures per file, the peak level and the average level over the whole clip. That is what makes this countable rather than arguable.
The set below is the 280 single-syllable clips at levels one and two, 70 per tone, 279 of them from one speaker. One speaker is the point rather than a limitation: with the voice held constant, tone is the only thing that varies.
One thing these figures are not: a measurement of pitch. Nothing in this manifest, and nothing in the checks that guard it, looks at the pitch of a clip or certifies that a tone is correct. The audio is trusted because of where it came from, and the numbers below are loudness readings of the file. Keep the two apart, because the whole argument here is that loudness and tone are not the same thing.
All four tones peak at the same level
The loudest moment of a first tone, a second tone, a third tone and a fourth tone are, for practical purposes, the same moment at the same level. Across 280 clips the four peak averages span 0.13 dB, which is far below anything a listener can detect and probably below what the recording chain could hold apart on purpose.
| Tone | Clips | Peak level, mean | Average level, mean | Duration, mean |
|---|---|---|---|---|
| First | 70 | -1.21 dBFS | -15.93 dBFS | 1.052 s |
| Second | 70 | -1.33 dBFS | -16.31 dBFS | 1.083 s |
| Third | 70 | -1.35 dBFS | -19.90 dBFS | 1.330 s |
| Fourth | 70 | -1.29 dBFS | -17.88 dBFS | 1.012 s |
Read the peak column first and then stop reading it, because it has nothing to say. Whatever made those four numbers agree to a tenth of a decibel, whether it is the speaker or the levelling in the recording chain, the consequence for you is the same: in this drill, turning your attention to how loud a clip is gives you nothing at all to work with.
The averages are not the same, and they run the wrong way
The average level over the whole clip does separate the four, and it separates them in the direction nobody expects. The fourth tone is the one every course describes as strong, abrupt, emphatic. It comes in at -17.88 dBFS, almost two decibels below the first tone, which nobody describes as anything in particular.
The third tone is lower still at -19.90, but the third tone has an obvious excuse: it is also the longest of the four by a wide margin, so the same peak is being averaged over more time. The fourth tone has no such excuse. It is the shortest of the four and still the quieter one.

The bowls are the shape of it. Every rim is at the same height, which is the peak, and every one holds a different amount, which is the average. Nothing about the rim tells you what is inside.
This is not the length of the clip talking
The obvious objection is that average level is just duration in disguise: pad any recording with more quiet and its average drops. It is a good objection and the corpus answers it twice.
- Inside each tone, the longer clips are slightly louder on average, not quieter. The correlation between duration and average level runs from +0.08 to +0.28 depending on the tone. Whatever is dragging the fourth-tone average down, it is not the fact that quiet time counts.
- Compare only clips of the same length and the gap survives. In the band from 0.888 to 1.100 seconds, where first, second and fourth tones all overlap, first tones average -15.97 dBFS and fourth tones -18.03. That is 2.06 dB between two piles of clips matched for duration. Not one third-tone clip is short enough to appear in that band at all.
So the difference is in the shape of the energy across the syllable rather than in how much of it there is. A fourth tone reaches full level early and spends the rest of its short life coming down. A first tone reaches full level and stays near it. That is why the peaks match and the averages do not, and it is a much better description of a fourth tone than loud ever was.
What English does with loudness instead
English is a stress language, and stress is carried by a bundle of things at once: a stressed syllable is longer, louder and higher than the ones around it, and the vowel in an unstressed one is reduced towards a schwa. You have spent your whole life reading that bundle as a single signal, and it tells you which syllable of a word matters and which word of a sentence matters. It never tells you which word it is.
Mandarin does the opposite. Pitch is spent on the identity of the syllable, not on its prominence, and loudness is left doing very little lexical work. When an English speaker meets a fourth tone and reports that it sounds angry, or forceful, or shouted, the report is real, and it is the loudness-and-emphasis circuit firing on something that is actually a pitch fall. The same misrouting sends a rising second tone to the part of the brain that handles questions. This is one concrete reason the tones are hard for English speakers, and it is a more useful reason than a general appeal to unfamiliarity: the ear is not failing to hear, it is filing what it hears in the wrong drawer.
The measurement above is the check on that. If the fourth tone really were the loud one, the English habit would work by accident and nobody would need retraining. It is not, so the habit produces a confident wrong answer instead of a useful one.
What to listen for instead
Three of the four cues people reach for first are duller than they look. Loudness is flat across the tones at the peak, and duration only separates the third tone. What is left is pitch direction, and that is the whole of it: level, rising, dipping, falling.
In practice that means one specific correction. If a clip strikes you as emphatic, do not record that as evidence for a fourth tone. Ask instead whether the pitch ended lower than it started. Those two questions feel like the same question and they are not, and separating them is most of what improving at the four tones consists of early on. Level one of Tingyin runs without an account, so the retraining costs nothing to start.
One honest limit. These are citation forms, syllables spoken alone and deliberately, in a curated drill corpus. Connected speech does other things to loudness, and a stressed word in a Mandarin sentence really is louder than the words around it. The claim here is narrower and it is the one that matters when you are staring at a drill: within a single syllable, loudness does not tell you which of the four tones you just heard.
Frequently asked questions
Is the fourth tone louder than the other Mandarin tones?
No, and in the recordings Tingyin plays it is quieter. Across 280 single-syllable clips the four tones peak within 0.13 dB of each other, and on average level the fourth tone sits about two decibels below the first, even when only clips of the same length are compared. What makes a fourth tone sound forceful to an English speaker is the fall in pitch and the early peak in energy, not the volume.
Why are Mandarin tones so hard for English speakers specifically?
Because English already uses pitch and loudness together, for a different job. In English a syllable that is longer, louder and higher is the stressed one, and the bundle marks emphasis rather than word identity, so a Mandarin pitch contour gets filed as attitude instead of as part of the word. Tingyin drills the four contours in isolation for that reason: the problem is a habit of interpretation, not an inability to hear.
How do I stop hearing the fourth tone as angry?
Replace the judgement with a question about direction. Instead of asking whether the syllable sounded forceful, ask whether it ended lower than it started, which is the only thing that actually distinguishes a fourth tone from a first. It takes a few hundred trials with immediate feedback before the second question becomes the automatic one, and that is what a drill is for.
What is the difference between peak level and average level in an audio clip?
Peak level is the single loudest instant in the file, and average level is the whole clip averaged over its duration. Two clips can share a peak and differ by several decibels in average, which happens whenever one of them spends more of its time well below its own peak. In this corpus the four tones share a peak almost exactly and differ in average, which is a statement about the shape of each syllable rather than its volume.
So am I imagining it when a fourth tone sounds like someone snapping at me?
Not imagining it, just mislabelling it. Something real is happening in the signal, and it is a fast fall from the top of the pitch range in a syllable that is over quickly. English marks finality and force the same way, so the sensation is genuine and the conclusion drawn from it is wrong, which is a much easier thing to fix than a hearing problem.