Should a Tone Drill Show You the Pinyin?
Not before you answer. A tone mark names the tone, so a drill that shows the pinyin while the clip plays is testing whether you can read a diacritic. It is also wrong about the tone often enough to matter: in Tingyin's corpus, 55 of the 357 multi-syllable recordings are words whose written tones are not the tones the speaker says.

A visible spelling changes which skill is being tested
Pinyin is not a hint. The mark above the vowel is the answer, written out: mā is first tone because the bar says so, and mǎ is third because the hook says so. Put that on screen while the audio plays and there is nothing left to work out. The learner reads, recognises a diacritic they already know, and clicks. Accuracy goes up, the session feels good, and nothing has been trained.
This is easy to miss because the failure looks like success. The thing that makes tone identification hard is that the four categories are not there yet, so pitch arrives as one continuous smear and the ear has nowhere to put it. Building the categories means making a forced choice from sound alone and finding out immediately whether it was right, which is the argument in what actually works in tone practice. A visible tone mark removes the forcing. There is no choice to make.
Which is not an argument for hiding the spelling altogether. It is an argument about order: answer, then see it.
And the spelling is not even reliable
There is a second reason, and it is the one that turns a training-design preference into a correctness problem. Mandarin rewrites tones in connected speech, and pinyin as it is normally written does not follow. The word is spelled with the tones the syllables have on their own, and the speaker says something else.
Counted in Tingyin's own corpus on 27 September 2026: of 637 human recordings, 357 carry more than one syllable, and in 55 of those 357 the written tones and the spoken tones disagree. That is about one word in seven. All 55 sit in levels 3 and 4, where the words are long enough for the rules to bite: 39 in level 3, 16 in level 4.
| Word | Written | Spoken | What moved |
|---|---|---|---|
| 水果 | shuǐ guǒ [3,3] | shuí guǒ [2,3] | Two third tones met |
| 不对 | bù duì [4,4] | bú duì [2,4] | 不 before a fourth tone |
| 一半 | yī bàn [1,4] | yí bàn [2,4] | 一 before a fourth tone |
| 大使馆 | dà shǐ guǎn | dà shí guǎn | Third tones met in the middle |
Three rules account for all 55. Twenty-five are two third tones meeting, where the first becomes a second. Eighteen are the number 一. Twelve are the negator 不. The mechanism behind each is in the tone sandhi rules; what matters here is the arithmetic. A learner reading tone marks off the screen gets the right answer for 302 of those 357 words and the wrong one for 55, and never finds out, because the drill agreed with them.
One number makes the point on its own. Twenty-three words in the corpus are written as two third tones in a row, and not one recording anywhere in the 637 is spoken as two third tones in a row. The sequence the spelling shows most often in that group is a sequence that does not occur.

How Tingyin orders it
Tingyin is a Mandarin tone-listening trainer: a clip plays, you pick the tone, and that is the whole exercise. It never asks you to say anything and never scores your voice. Every clip is a human recording carrying its source, speaker, licence and checksum, 637 of them, and level 1 needs no account at all.
While a clip is playing there is no writing on the screen. No characters, no pinyin, no gloss, nothing but the play control and the four tone buttons. The word appears the moment the answer lands, characters and pinyin together, and on a word where the two forms diverge a panel shows both, labelled Written and Spoken, with a line saying you were graded on what was spoken. The disagreement is the teaching moment, so it is shown deliberately rather than smoothed over. It just arrives second.
If you are building your own practice
Most people drilling tones are not using a purpose-built tool. They are using flashcards, a podcast, or a list of words with audio, and in all three the spelling tends to be visible by default. Four changes are worth making.
- Put the audio on the front of the card and the pinyin on the back. The common setup is the reverse, which makes a card that tests reading.
- Commit to an answer out loud or with a keypress before you reveal. Thinking "probably second" without committing is how a wrong guess gets quietly upgraded once the answer appears.
- Keep the characters hidden too, at least at first. If you know the word, the character tells you the tone as surely as the diacritic does.
- Expect the sandhi words to feel like errors. When a word spelled with two third tones sounds like second-then-third, your ear is right and the page is behind. Mark it correct.
Tone numbers have the same problem as marks, incidentally, and for the same reason: ma3 is an answer in transport format. The difference between the two notations is set out in the piece on tone marks, and neither belongs on screen before the guess.
Frequently asked questions
What is the best way to drill Mandarin tones?
Audio first, forced choice, immediate feedback, and the spelling revealed only after you have committed. Tingyin is built to that order: nothing is written on the screen while the clip plays, and the characters and pinyin appear with the verdict. Flashcard decks can be set up the same way by putting the recording on the front of the card, which is the opposite of how most shared Mandarin decks arrive.
How do I stop reading the tone instead of hearing it?
Remove the text from the moment of decision, because as long as it is there the eye will win. That means no pinyin, no tone numbers and, if you know the word, no characters either. Anki and similar tools support an audio-only front; the change takes a minute in the card template and costs nothing else.
Is it safe to trust the pinyin for the tone of a word?
For a single syllable in isolation, yes. For a word of two or more syllables, not reliably: 55 of the 357 multi-syllable recordings in Tingyin's corpus are spoken with tones other than the ones they are written with, roughly one in seven. Pinyin is normally written in citation tones, so the page shows you what the syllables do alone rather than what they do together.
What is the difference between the written tone and the spoken tone?
The written tone is the tone a syllable carries on its own; the spoken tone is what comes out when it stands next to other syllables. Two third tones in a row are the clearest case: 23 words in the corpus are written that way and not one recording in all 637 is spoken that way, because the first third tone always becomes a second. The number 一 and the negator 不 shift in the same fashion, and between them the three rules account for every disagreement in the corpus.
I can get them right in the app but not when someone talks to me. Is the app the problem?
Possibly, and the first thing to check is whether anything was on screen when you chose. Getting tones right with the pinyin visible is a different task from getting them right from sound, and it transfers poorly for exactly that reason. If the text was hidden and it still does not transfer, the likelier culprits are speed and the number of voices you have heard, which is a different problem with a different fix.