The Second Tone Is the One That Holds Still
Across the 1,127 syllables Tingyin ships, the second tone is the one whose share of what you hear barely moves: 25% in the single-syllable drills, 24.7% at level three, 23.3% at level four. Every other tone shifts hard over that range. The fourth climbs from 25% to 35.4%, the third falls to 15.5%, and a neutral tone that does not appear in the drills at all takes 9% of level four. So the second tone is the one worth being solid on before you leave the drills, because it is the only one whose weight in real words is the weight the drills gave it.

The tone nobody writes a page about
The third tone gets the attention, because it is the one that misbehaves: it is usually not the dip the textbook draws, and it changes into something else when another third tone follows it. The neutral tone gets attention because it has no shape at all. The fourth is short and unmistakable, and the first is a flat line.
The second sits in the middle of all that and gets described in one sentence, usually the wrong one. It is not the English question rise, and it does not start at the bottom of the voice. What it does is start around the middle of the speaker's range and climb from there, and the distance it climbs is smaller than the diagram suggests.
Where these numbers come from
Everything below is counted from the audio manifest that ships with the app: 637 clips, 1,127 syllables, each entry recording the word, the tones as they are actually pronounced, the level it appears at, the speaker, the licence and a checksum. It is a count of one corpus rather than a claim about Mandarin at large, and one corpus is a small thing. It is, however, the corpus a learner using this app will actually hear, which makes it the right one for the question being asked here.
Levels one and two are single syllables and are balanced on purpose: 20 clips per tone at level one and 50 per tone at level two, which is 70 apiece across the 280 of them. Levels three and four are real multi-syllable words, and nobody balanced those, because words are not distributed the way a drill is.
What shifts when the drills stop
| Tone | Levels 1 to 2 | Level 3 | Level 4 |
|---|---|---|---|
| First | 25.0% | 22.8% | 16.6% |
| Second | 25.0% | 24.7% | 23.3% |
| Third | 25.0% | 19.4% | 15.5% |
| Fourth | 25.0% | 23.0% | 35.4% |
| Neutral | 0% | 10.1% | 9.1% |
Read down the second row and then down any other. The second tone gives up 1.7 points across the whole range. The fourth gains more than ten, the third loses nearly ten, and the neutral tone arrives from nothing. Four of the five rows describe a different listening problem at level four than they described at level one.
This has a practical consequence that is easy to misread. If your accuracy moves when you change level, part of that movement is the distribution rather than your ear. You are being asked a different set of questions, weighted differently, so a score is not comparable across the boundary. The one row you can compare is the second.

It is also not the long one
A second cue people reach for is length, and here the second tone is unremarkable, which is itself useful. Measured over the 280 single-syllable clips, 70 per tone and almost all one speaker, the mean durations are 1.05 seconds for the first tone, 1.08 for the second, 1.33 for the third and 1.01 for the fourth.
- The third tone stands nearly a third of a second clear of everything else. Length identifies it.
- The first, second and fourth sit inside 70 milliseconds of each other. Length tells you nothing about which of the three you heard.
So the length cue that rescues you on the third tone is exactly no help on the second, and a learner who has been leaning on duration without noticing will find it quietly stops working. What 280 recordings show about tone length goes through the measurement in full.
Some of the second tones you hear are third tones
This is the part that makes the second tone genuinely slippery rather than merely undramatic. When one third tone is followed by another, the first is pronounced as a second tone. 可以 is written with two third tones and said ké yǐ. 所以 and 手表 are the same. The corpus records tones as they are spoken, and the evidence is clean: across 357 multi-syllable clips there is not a single adjacent third-third pair, and there are 33 second-third pairs, some of which are those words.
For a listener this means the question is genuinely ambiguous at the level of sound. You can hear a first syllable correctly as a second tone and still be looking at a word spelled with a third, and no amount of ear training resolves that, because there is nothing there to resolve. What resolves it is knowing the word. Why 3+3 sounds like 2+3 covers the rule and the two other words that shift.
What to do with this
Tingyin is a Mandarin tone-listening trainer: it plays a clip and asks which tone you heard, and it does nothing else. Level one needs no account and levels one and two are free. Every clip is a human recording carrying its source, speaker, licence and checksum, which is why the numbers above exist to be counted at all.
- Get the second tone solid during the drills. It is the only tone whose share of your listening does not change when you leave them, so the work transfers at full value.
- Expect the second-versus-third confusion to be the last one to go, and do not treat every instance of it as a failure of your ear. Some of them are sandhi.
- When your score moves after a level change, check the distribution before you conclude anything about yourself. The fourth tone is ten points more common at level four.
There is a broader version of this argument in why Mandarin tones are hard to hear rather than to say, and the pairwise view in why two syllables are the wall.
Frequently asked questions
What does the second tone in Mandarin sound like?
A rise that starts around the middle of the speaker's pitch range and climbs from there, rather than one that starts at the bottom. It is often compared to an English question rise, which misleads: the English rise sits at the end of a whole sentence and the Mandarin one lives inside a single syllable, so it is faster and smaller than the comparison implies. Tingyin plays human recordings and asks which tone you heard, which is a different exercise from reading a description of the shape.
Why do I confuse the second and third tones?
Because they share their lower half. Both spend time near the bottom of the range and differ in what happens afterwards, and in connected speech the part that differs is short. There is also a real ambiguity underneath the perceptual one: a third tone followed by another third tone is pronounced as a second, so some of what you hear as a second tone is written as a third. That part is not an error in your ear.
Is the second tone longer than the others?
No. Across the 280 single-syllable clips in Tingyin's corpus the second tone averages 1.08 seconds against 1.05 for the first and 1.01 for the fourth, which is inside the range of ordinary variation. The third tone is the outlier at 1.33 seconds. Length is a usable cue for spotting a third tone and no help at all for spotting a second.
Which Mandarin tone is the most common?
In this corpus the fourth, once you get past the drills: it takes 35.4% of the syllables at level four against 23.3% for the second and 15.5% for the third. The single-syllable levels are balanced at 25% each by construction, so a learner who has only done drills has heard a distribution that does not resemble the words. A larger corpus would give different percentages, and the direction of the shift is the part worth carrying.
Should I drill the second tone on its own or in pairs?
On its own first, because a single syllable is where the shape is clearest and where the second tone is the least contaminated by what is next to it. Move to pairs once single syllables are reliable, since that is where sandhi and the neutral tone start to appear and where most remaining errors live. Tingyin's levels are built in that order for this reason: single syllables at levels one and two, real words at three and four.