Why a Tone Drill Should Not Be Perfectly Random
A perfectly random tone drill is a worse drill. Shuffle four tones properly and long runs of the same answer are not a rare accident - in an evenly balanced hundred-item session you should expect about four stretches of three or more identical tones, and every one of them quietly stops being a listening test. Tingyin's drill order is random but capped: the same tone never asks three times in a row unless the level's own contents make that arithmetically impossible. The order is fixed when the level loads, so a reload does not deal you a new hand mid-session.

What a run of the same tone does to you
Answer second tone correctly, and the drill plays another second tone, and another. By the third one you are no longer identifying anything. You have noticed a pattern and you are riding it, and the fourth item gets answered before the clip has finished playing.
This is not laziness, it is how the ear works. Repeated exposure to one category without contrast is exactly the condition under which discrimination gets worse rather than better - the whole point of a forced-choice drill is that the next answer is genuinely open. A run closes it. And the damage is invisible from the score: you got those four right, so the session reports a good run when what it actually recorded was one identification and three guesses that happened to land.
Runs are not rare, which is the surprising part
People assume a shuffle protects them from this. It does not. Any given position has a one-in-sixteen chance of being the third of three matching tones, and over a hundred items those chances accumulate into something you notice. Simulating the real case - an honest shuffle of a level holding twenty-five clips of each tone - the average session contains about four separate stretches of three or more. Not a freak occurrence: a regular feature.
A hundred-item session is not unusual either. So the untreated version of this drill hands almost every learner several stretches where they are being tested on their pattern recognition instead of their ears, and neither the learner nor the score can tell which is which.

Random, then constrained
The fix is not to make the order less random in any way you would feel. Tingyin shuffles the level normally, then deals the items out under one rule: the same tone may not be asked more than twice in a row. Where several tones are available for the next slot, the one with the most items still waiting goes first.
That last detail is the part that makes it work rather than merely sound good. Dealing the most-frequent tone first is what keeps the arrangement from painting itself into a corner - the alternative, picking at random from whatever is legal, reliably strands forty items of one tone at the end of the queue with nothing left to interleave them with. Most-frequent- first only fails when the level's contents make failure unavoidable, and then it stretches one run rather than giving up.
| Ordering | Runs of three or more | What the learner does |
|---|---|---|
| No shuffle - grouped by tone | Constant | Stops listening within the first group |
| Plain shuffle | About four per hundred items | Coasts through each run without noticing |
| Run-capped shuffle | None, unless the level forces one | Has to identify every item |
The order is dealt once, not re-dealt on every render
A smaller decision with a visible consequence: the order is settled when the level's clips are fetched, not while the screen is drawing. So the first item the page shows you is the first item it meant to show you, and a refresh part-way through does not reshuffle the remainder into a fresh sequence. Restarting or replaying deals a new order deliberately; nothing else does.
Tingyin is a Mandarin tone-listening trainer and it does only the one thing: a clip plays, you pick the tone. Every clip in it is a human recording that carries its source, speaker, licence and checksum - 637 of them, 622 under CC BY-SA 4.0 and 15 under CC0. Level 1 runs with no account at all and level 2 is free once you sign in. Nothing about the ordering rule changes with a subscription; it is how the drill is built.
What this does not fix
- It does not make the drill harder in the way that matters most. Interleaving tones is a floor, not a curriculum - which pairs and which speakers you meet still decide how much you learn.
- It says nothing about your production. Nothing here judges how you sound, and what a listening drill can and cannot train goes through that boundary properly.
- Two in a row still happens, on purpose. Banning repeats entirely would make the drill predictable in the opposite direction - the one thing you could always rule out is the answer you just gave.
If you are earlier than this and the four contours themselves are still the problem, why the tones are hard to hear rather than to say is the place to start; if the single tones are solid and words are still slipping, the two-syllable pairs are where the wall actually is.
Frequently asked questions
What is the best order to drill Mandarin tones in?
Interleaved rather than blocked - mix the tones instead of doing thirty first tones and then thirty second tones. Tingyin deals its drill randomly but caps the same tone at two in a row. Anki does the opposite by default: its queue groups a new card with its own repeats, which is right for vocabulary and wrong for tone discrimination.
How do I stop guessing my way through a tone drill?
Watch for the moment you answer before the clip has finished. That is the signature of riding a pattern rather than listening, and it usually happens inside a run of the same tone. If your drill lets you group tones together, stop doing that first.
Is it safe to trust a drill score as a measure of progress?
Only if the order is interleaved. A score from a blocked drill mostly measures how quickly you spotted the block. Tingyin's ordering rule exists so that its own scores mean something, and the honest limit is that no listening score says anything about how you sound.
What is the difference between shuffling and interleaving?
Shuffling is random and interleaving is a constraint on the randomness. A shuffle can legitimately hand you five second tones in a row; an interleaved order will not, because avoiding that is the thing it is for. Most apps do the first and describe it as the second.
Why not just ban the same tone twice in a row entirely?
Because that leaks the answer. If a repeat were impossible, then after every item you could rule one of the four tones out for free, and a quarter of the difficulty would disappear. Two in a row is allowed precisely so that the previous answer tells you nothing.