How Long a Tone Drill Should Actually Run
Somewhere between five and fifteen minutes, most days, and the top of that range is generous. Sustained-attention performance in the general literature starts sliding inside the first quarter of an hour, and distributed practice beats one long block on almost everything anyone has tested it on. The exact minute your own returns drop is individual and nobody can hand it to you, but it is measurable and there are four reliable signs you have already passed it.

The short answer, and why it is a range
A tone drill is a sustained forced-choice task. You listen to something short, you commit to one of four categories, you get told whether you were right, and then it happens again. That is one of the most heavily studied shapes of task in psychology, and the finding that comes back from it over and over is that performance on this kind of thing does not hold steady. It declines with time on task, in a way people do not notice while it is happening.
So the useful part of a session is bounded, and the boundary is much earlier than the point where you feel tired. What nobody can tell you is where yours is. It moves with sleep, time of day, how noisy the room is, and whether you are working on a contrast you have already solved or on the one that is still costing you. Anyone quoting a single number for this is quoting a number they made up.
The range worth starting from is five to fifteen minutes. Below five you have not done enough trials for the session to mean anything. Above fifteen you are usually paying full attention costs for trials that are no longer teaching you anything.
What degrades, and how fast
The effect has a name and a long history. Norman Mackworth's 1948 experiments on radar-style monitoring found that detection performance fell off within the first half hour of watching, and the decline has since been reproduced across a wide range of sustained attention tasks, often with a measurable drop inside the first fifteen minutes. The specifics vary a great deal with the task; the shape does not.
Worth being precise about what is degrading, because two different things get called fatigue and only one of them applies here.
- Your ear is not the thing that gets tired. Genuine auditory fatigue, in the sense of a temporary shift in hearing threshold, needs sound levels well above anything a listening drill uses. If you are drilling at a sensible volume, your cochlea is fine an hour in. If you are drilling at a volume where it would not be, you have a different problem, and what headphones do and do not fix is the place to start.
- Attention is the thing that gets tired. Specifically the part that holds a perceptual decision to a criterion, trial after trial, with no external structure forcing you to. That is the resource the vigilance literature measures, and it is the one running your drill.
- The decline shows up as a change in strategy, not effort. People do not stop trying. They start deciding differently: slower, more deliberately, with more reasoning and less listening. Accuracy can stay flat for a while during this, which is why it is easy to miss.
That last point is the one that matters for a tone drill in particular, because the skill you are trying to build is a fast perceptual judgement. A trial you solved by thinking hard about it is not practice at the thing you want. It is practice at a different, slower task that happens to have the same answer.
Why short and frequent beats long and rare
Two separate reasons stack, and they point the same way.
The first is the spacing effect. Splitting a fixed amount of practice across several sessions produces better retention than the same amount in one block. This is one of the oldest and most reproduced findings in the study of learning, and it holds across material types and across intervals. Thirty minutes as three tens beats thirty minutes as one thirty, before you account for attention at all.
The second is specific to perceptual learning. In auditory discrimination training, improvement frequently shows up between sessions rather than during them. Learners finish a session no better than they started and come back the next day measurably better, and consolidation during the interval, including during sleep, is the usual account. If that is what is happening, sessions are not just containers for practice. The gap between them is doing work of its own, and a schedule with no gaps has removed the part that pays.
This is the same argument as the one behind the daily streak counter, from the other end. That article is about turning up at all, which is the harder problem. This one assumes you turned up and asks what to do with the next ten minutes.

What a good session looks like
Numbers help here, so these use a real corpus. Tingyin is a Mandarin tone-listening trainer: a clip plays, you pick one of the tones, it tells you immediately. Level 1 is 80 clips, 20 syllables recorded in all four tones, and the audio in it adds up to 91 seconds. Level 2 is 200 clips and 222 seconds of audio. Every one of the 637 clips across the whole corpus is a human recording carrying its source, speaker, licence and a sha256, and levels 1 and 2 are free, with level 1 running in a browser with no account.
The audio is a small fraction of the time. What sets the length of a session is how long you spend deciding.
| Session | Items | Roughly | What it is good for |
|---|---|---|---|
| Half a level 1 run | 40 | 3 to 4 minutes | A day you would otherwise skip entirely |
| A full level 1 run | 80 | 6 to 8 minutes | The default daily session |
| Half a level 2 run | 100 | 8 to 12 minutes | A longer session while attention is still good |
| A full level 2 run in one sitting | 200 | 17 to 23 minutes | Past the useful window for most people |
| Three short runs across a day | 3 by 40 | 12 minutes total | The best return of the five, and the hardest to arrange |
The shape inside the session matters as much as the length. A minute of a contrast you already have gets your ear anchored to the speakers and the volume, and it is not wasted. Then the pair that is actually costing you, which is where nearly all the remaining gain lives once the curve has flattened. Then stop, on purpose, while it is still going well.
Stopping while it is going well is a motivational point rather than a claim about learning. A session that ends in a run of errors changes whether you open the thing tomorrow, and tomorrow is where the gain from today shows up.
Four signs you have gone past useful
None of these is about feeling tired. They all appear well before that, and they are specific enough to notice in the moment.
- You start replaying clips. The first time you reach for a second listen on an item you would have answered immediately ten minutes ago, the session has turned. A replay is not cheating, it is a symptom.
- You start narrating. If you catch yourself thinking "that went down and then came back up, so it must be", you have left perception and entered reasoning. The reasoning gets the answer right often enough to hide the fact that you are no longer training the skill you came for.
- Your errors spread out. A learner with a real perceptual gap gets things wrong in a pattern: two cells of the confusion grid hold most of it. When the errors start landing everywhere, including on contrasts you find easy, that is attention rather than perception, and no amount of it teaches you anything.
- You stop reading the feedback. The moment where you press, see the result and move on without registering it is the moment the session became button-pressing. This is the one people are most reluctant to admit, because the counter keeps going up.

How to find your own number
The clean experiment is to run the same block of items at the start and the end of a session and compare. Almost no drill lets you construct that, and Tingyin does not: its order is random within a level and fixed when the level loads, so the second half of a run is different material from the first. Comparing halves inside one session mostly measures which items you happened to get.
What works instead is cruder and takes a week.
- Note the minute, not the score. Each session, write down the clock time at which the first of the four signs appears. That is one observation. Six of them, across six days, will cluster more tightly than you expect.
- Compare across days, not within one. Run the same level for a week at a fixed length and watch the accuracy for that level move. If it is climbing, the length is fine. If it is flat while you are still finding the material hard, try shorter sessions before you try longer ones.
- Change one thing at a time. Length, time of day and which contrast you are working on all move the result. Changing two at once tells you nothing, and a week is short enough that you will be tempted to.
And a caution about the whole exercise. This is worth doing once, to find out roughly where your own edge sits, and it is not worth doing continuously. A learner who spends every session monitoring their attention has added a second sustained-attention task on top of the first one. Find the number, set a timer for it, and then go back to listening.
Frequently asked questions
What is the best length for a Mandarin tone practice session?
Five to fifteen minutes, most days, with the useful part usually nearer the bottom of that range than the top. Sustained-attention performance in the general literature begins sliding within the first quarter of an hour, and distributed practice beats one long block, so three short sessions across a day outperform one session of the same total length. A full run of Tingyin's 80-clip first level takes six to eight minutes at a steady pace, which makes it a reasonable unit for a daily session.
How do I know when to stop a listening drill?
Stop at the first of four signs rather than at a fixed score: you replay a clip you would have answered immediately earlier, you start explaining the contour to yourself in words, your errors spread across contrasts you normally find easy, or you notice you are no longer reading the feedback. Each of those means the session has turned into a slower task with the same answers. None of them feels like tiredness, which is why a timer set from your own observations works better than waiting to feel done.
What is the difference between mental fatigue and ear fatigue in tone practice?
Ear fatigue in the strict sense is a temporary shift in hearing threshold and needs sound levels far above anything a listening drill uses, so at a sensible volume it is not what is happening to you. What runs out is the attentional resource that holds a perceptual decision to a consistent criterion trial after trial, and that one is measurable within minutes. The practical difference is that resting your ears does nothing and changing task does, which is why a ten-minute drill twice a day works and a twenty-minute drill once does not.
Is it safe to practise Mandarin tones for an hour a day?
Safe, yes, at a reasonable listening volume, and largely wasted as a single block. Roughly the last forty minutes of it will be spent in the deliberative mode that trains a slower skill than the one you want, and the additional trials do not consolidate the way earlier ones do. If you genuinely have an hour, split it: a short tone drill, then vocabulary in Anki or Pleco, then another short drill later. The tone perception part is the one with the low ceiling per session.
Why do I get worse at tones the longer I practise?
Because the way you are answering changes before your accuracy does. Deep into a session people stop hearing the contour and start reasoning about it, which is slower, more effortful and worse on exactly the borderline items that decide your score. The errors also stop being informative, since they scatter across contrasts you have already solved instead of concentrating in the pair you are actually working on. Coming back the next day usually shows a higher score than the end of the previous session did, which is the clearest evidence that the drop was attention rather than skill.