Is One Mandarin Tone Quieter Than the Others?
One of them is, and it is the third. Measured across the 280 single-syllable recordings Tingyin plays at levels one and two, third tones average 19.9 dBFS below full scale against 15.9 for first tones, a gap of four decibels. Every one of those clips peaks at the same level, within a seventh of a decibel. The third tone is not spoken more softly; it spends more of its length quiet.

Where these numbers come from
Tingyin is a Mandarin tone-listening trainer: a clip plays, you pick the tone, and that is the whole exercise. Every clip it ships is a human recording carrying its source, its speaker, its licence and a hash of the audio, so the corpus can be counted rather than estimated, and the manifest also records the duration and the level of each file. That is what makes a question like this answerable at all.
The set here is the same one behind the tone-duration measurements: the 280 single-syllable clips, exactly 70 per tone, all at levels one and two, which is the free part of the app and where level one needs no account. 279 of the 280 come from one speaker. For this particular question that matters more than usual, because comparing loudness across speakers is mostly comparing how close each of them sat to a microphone.
Two numbers per clip are worth separating. The peak level is the single loudest instant in the file. The average level is the energy across the whole file. They answer different questions, and here they disagree.
The third tone sits four decibels down
| Tone | Average level | Median | Range | Peak |
|---|---|---|---|---|
| First | -15.9 dBFS | -16.0 | -19.5 to -13.0 | -1.21 dBFS |
| Second | -16.3 dBFS | -16.2 | -19.4 to -11.7 | -1.33 dBFS |
| Fourth | -17.9 dBFS | -17.9 | -22.9 to -14.9 | -1.29 dBFS |
| Third | -19.9 dBFS | -19.8 | -27.6 to -14.6 | -1.35 dBFS |
The order is stable rather than marginal. Not one of the 70 first-tone clips sits below the median third-tone clip, and only five of the 70 third tones climb above the median first tone. The first and second tones are 0.4 dB apart, which is nothing; the fourth sits two decibels below them; the third is four decibels below the first and, in this corpus, is reliably the quietest thing in the drill.
The third tone is also the widest. Its clips span thirteen decibels between the quietest and the loudest, against six and a half for the first tone. The tone that is quietest on average is also the least consistent, which is a bad combination for a listener working through a set at a fixed volume.
The peaks are identical, and that is the whole story
Look at the last column again. All four tones peak within 0.14 dB of each other. The loudest instant of a third tone is exactly as loud as the loudest instant of a first tone, and if loudness were a property of the tone itself that could not be true.

What is actually happening is the shape. A third tone falls to the bottom of the speaker's range, sits there, and comes back. The bottom of the range is where the voice is at its weakest and often creakiest, and the clip is a third of a second longer than a fourth tone, so the quiet middle is both quieter and longer. A first tone holds a steady pitch for its whole length and has no such trough. Average the two files and the third tone loses; take the loudest moment of each and they tie.
This is worth saying carefully, because the same numbers support a false version. The third tone is not spoken more softly, and a speaker is not putting less effort into it. It has a low, weak stretch in the middle of it, and averaging over the file turns that into a number that looks like volume.
Why loudness is still not something to listen for
A four-decibel difference measured within one voice at one distance from one microphone is not a cue you can carry into a conversation. Turn the volume knob and every number above moves together. Move the speaker a metre further away and they move again. Put a second speaker in the room and the between-speaker variation swamps the between-tone variation entirely.
So the practical value of this measurement is not that you should listen for a quiet tone. It is that you should stop treating one particular failure as an ear problem:
- If three tones are comfortable and the third keeps slipping away, the first thing to check is the listening conditions rather than your ear.
- Anything that squeezes the dynamic range, including a phone loudspeaker in a noisy street, will take the middle out of a third tone before it touches the other three.
- Raising the volume raises everything, so it makes a quiet third tone audible without making it any easier to tell apart from a second tone. Those two confuse for reasons of pitch direction, not level.
What to change when one tone keeps disappearing
In order of how much difference it makes. Use headphones or earbuds rather than a laptop or phone speaker, because the fundamental frequency a tone lives on sits below what a small speaker reproduces at all. Practise somewhere quiet, since a low weak stretch is the first thing background noise covers. And be suspicious of a compressed or narrow-band source: what a phone line does to a tone explains why a voice call is a harder listening test than the drill is.
Then go back to the drill and treat a missed third tone as information about the tone rather than about you. Level one of Tingyin plays all four in the same set without an account, and the third tone is the one worth being slowest and most deliberate about.
Frequently asked questions
Which Mandarin tone is the quietest?
The third, measured across the 280 single-syllable clips Tingyin ships at levels one and two. Third tones average -19.9 dBFS against -15.9 for first tones, -16.3 for second and -17.9 for fourth, and not one first-tone clip in the set sits below the median third tone. The reason is the shape rather than the effort: a third tone dips to the bottom of the speaker's range and stays there, and averaging that trough over the file pulls the number down.
How do I stop losing the third tone when I am listening to Mandarin?
Fix the listening conditions before you blame your ear. Headphones instead of a laptop speaker, a quiet room, and a source that has not been compressed down a phone line will each recover more of a third tone than extra practice will, because what goes missing is a genuinely low and quiet stretch in the middle of the syllable. Tingyin plays the same clips at level one for free and without an account, so it is easy to test whether a change of hardware fixes it.
Is it safe to identify a tone by how loud it sounds?
No, and the same measurement that finds the difference is what rules it out. Four decibels is a real gap inside one voice at one distance from one microphone, and it disappears the moment the speaker, the room or the volume changes, none of which a listener controls in conversation. Use it as an explanation for why one tone keeps vanishing, never as a way to answer which tone you just heard.
What is the difference between the peak level and the average level of a clip?
The peak is the single loudest instant in the file and the average is the energy spread across the whole of it. In this corpus the four tones peak within 0.14 dB of each other while their averages spread over four decibels, which is exactly what you would expect from a tone that reaches the same maximum as the others but spends part of its length near the bottom of the voice. Duration feeds into it too: a third tone runs about a third of a second longer than a fourth.
My third tones sound fine in the app and vanish when a real person says them. Why?
Because a drill clip is a clean recording of one syllable and a conversation is not. Running speech shortens the tone, a room adds noise over exactly the quiet part of it, and a speaker halfway across a table is another ten decibels down before anything else happens. Tingyin trains the recognition and cannot supply the conditions, so the useful next step is deliberately harder listening rather than more of the same drill.