Why a New Voice Makes Mandarin Tones Hard Again
Two things change at level three, not one. The words go from one syllable to two, and the voice changes: of the 280 clips at levels one and two, 279 are the same speaker, and that speaker appears zero times at levels three and four. So an accuracy drop there has two possible causes, they need different practice, and the score cannot tell you which one you are looking at.

What actually changes at level three
Tingyin ships 637 recordings, and the manifest that describes them records the speaker of every one. Counting them by level produces a split the level descriptions do not mention and which changes how a drop in accuracy at level three should be read.
| Level | Clips | Syllables | Main voice |
|---|---|---|---|
| 1 | 80 | One | Chen Wang, all 80 |
| 2 | 200 | One | Chen Wang, 199 of 200 |
| 3 | 237 | Two | Yue Tan, 225 of 237 |
| 4 | 120 | Three or four | Yue Tan, 104 of 120 |
The two voices never meet. Yue Tan has no clips at all at levels one and two; Chen Wang has none at levels three and four. A learner working through in order therefore hears one person say 279 single syllables, and then, on the first item of level three, hears a different person say a two-syllable word. Both variables move on the same click.
A new voice is a real difficulty, not an excuse
A Mandarin tone is a shape inside one speaker's range rather than a pitch you could name in hertz. A high level tone from a low male voice can sit below a third tone from a high female voice in absolute terms and still be unmistakably the high one, because what identifies it is where it sits relative to the rest of that person's speech. Why tones are not musical is the same fact from the other side: perfect pitch buys you nothing here.
The consequence for training is the part that gets missed. A category learned from one voice is partly a category about that voice. You are not only learning what a second tone is; for a while you are learning what a second tone sounds like when Chen Wang says it. That generalises, and it generalises with exposure rather than instantly, so the first stretch with an unfamiliar speaker is genuinely harder and is supposed to be.

Two variables, one number
Level three is usually described as the point where two-syllable words arrive, and that is the harder problem on paper: four tones and a neutral make twenty combinations rather than four, and the tone pair grid is where most people actually stall. What the manifest adds is that the pair grid is not the only thing that changed, so attributing the whole drop to it is a guess.
The clips get longer too, which cuts the other way: level one and two clips average around 1.1 seconds, level three 1.24 and level four 1.46. More audio is usually easier, not harder. So the honest reading of a fall in accuracy at level three is that at least two things moved and one of them helped, and no single accuracy number separates them.
How to pull the two apart
The useful move is to change one variable at a time, which the corpus can just about support.
- Go back and re-drill level two for a session. Same voice, one syllable. If that is still comfortable, the tones themselves are intact and the problem is ahead of you rather than behind.
- Sit in level three long enough to stop noticing the voice. A single unhurried sitting is usually enough to stop hearing a stranger and start hearing tones, and the improvement across that first session is mostly this rather than anything about pairs.
- Then judge the pairs. Whatever is still wrong after the voice has become familiar is the real pair problem, and the plateau arithmetic explains why two unseparated tones cap a score however good the rest is.
What this corpus cannot give you
Eight speakers, and two of them carry 608 of the 637 clips. The remaining six contribute 29 recordings between them, from a community source, and they appear scattered through levels two to four rather than as a block you could drill. That is enough to make the point that tones survive a change of speaker and not enough to train the generalisation properly.
So the last step is not in the app. Once level three is comfortable, listen to Mandarin from somebody neither of these two speakers, at a length nobody drills: a podcast, a phone call, a person in front of you. Everything the corpus can teach about provenance, including what it does not record, is in what a tone corpus records about accent. The clips are human recordings with a named speaker, a licence and a checksum each, which is what makes a count like this one possible at all, and it is also the ceiling: 637 clips is a training set, not a language.
Frequently asked questions
Why do Mandarin tones sound different from different speakers?
Because a tone is a movement inside the speaker's own pitch range rather than a fixed frequency, so the same second tone arrives at completely different hertz from a low voice and a high one. Listeners normalise for this automatically once they have heard enough of a person, which is why an unfamiliar voice is briefly harder and then stops being harder. Tingyin ships 637 human recordings from eight speakers, so the change is audible in the drill itself rather than something to take on trust.
Is level three harder because of the words or because of the voice?
Both change at once, which is exactly the problem. Level three moves from one syllable to two and also moves from a speaker who records 279 of the first 280 clips to one who records none of them, so a single accuracy figure cannot tell you which change cost you the points. Drilling level two for a session isolates the tones from the pairs, and the new voice needs a full sitting of its own before the comparison is fair.
How many voices do I need to hear to learn Mandarin tones?
More than one, and the second one does most of the work, because it is the first evidence that the category is about the tone rather than about a person. Beyond that the returns come from variety of speaking style and speed rather than from a headcount. Tingyin's corpus has eight speakers with two carrying 608 of the 637 clips, so it can demonstrate the effect and not train it to completion.
What is the difference between a tone pair problem and a speaker problem?
A pair problem is stable and specific: the same two tones are confused in the same direction, session after session, and it does not improve just by listening more. A speaker problem fades, usually inside one sitting, and it affects everything rather than one pair. If accuracy comes back up across the board within a session it was the voice; if two particular tones stay tangled it was never the voice.
I was fine and then suddenly I could not hear anything. Have I gone backwards?
Almost certainly not, and the shape of the drop is the giveaway: real loss is gradual and this is a cliff at one specific point in the sequence. What happened is that the material changed underneath you, in two ways at the same time, on a single step. Tingyin keeps levels one and two free, with level one needing no account at all, so going back to re-check the easy case costs nothing but the sitting.