Can Speech Recognition Tell If Your Tones Are Right?
Not the way you are using it. Dictate a sentence and the recogniser repairs your tones out of the surrounding words and hands you the characters you meant, which reads as a pass and is not one. Dictate isolated syllables and the same recogniser becomes a genuinely harsh tone test, because there is nothing left for it to repair with.

Why everybody tries this
It is the obvious idea and it is nearly a good one. You want to know whether your tones are right, you do not have a tutor in the room, and there is a machine in your pocket that converts speech to Chinese characters for free and writes down what it thought you said. Unlike a person it has no manners, so it should have no reason to be kind about it.
The trouble is that politeness was never the mechanism. A speech recogniser flatters you for a reason that has nothing to do with wanting to, and understanding that reason is what turns the tool from useless into sharp.
A recogniser is two things, and only one of them is listening
Every system that turns speech into Chinese text is doing two jobs at once. One half listens to the sound and proposes what it might have been. The other half knows what sequences of words are likely in Chinese and picks the reading that makes sense. The second half is the reason dictation works at all, and it is also the reason it cannot grade you.
Mandarin gives that second half an enormous amount to work with. There are 409 syllables before you count tone, which is a small inventory for a language with tens of thousands of words, so the language is dense with syllables that sound alike and are told apart by tone, by context, or by both. Say a wrong tone inside a sentence and the surrounding words still point at one answer. The machine takes it, writes the characters you meant, and never mentions that the pitch it heard did not match.
Tingyin is a Mandarin tone-listening trainer, built the other way round: a clip plays, you pick which of the four tones you heard, and you are told immediately whether you were right. Levels 1 and 2 are free and level 1 needs no account at all. It exists because the forced choice is the part that dictation cannot give you.

Take the context away and it starts telling the truth
The fix follows directly from the mechanism. The repair runs on context, so the test has to remove the context. Say one syllable, alone, with nothing before or after it. The half that knows what Chinese sentences look like has nothing to condition on, the acoustic half is left holding the decision by itself, and what comes back on the screen is much closer to what you actually produced.
It is not a laboratory instrument and it should not be described as one. A recogniser can still guess, it is tuned for connected speech rather than for single syllables, and a stubbornly wrong character sometimes means the vowel rather than the tone. What it does reliably is stop agreeing with you. That alone puts it far ahead of conversation, which fails in the direction that flatters you every single time.
| What you dictate | What the language half can do | What the result tells you |
|---|---|---|
| A whole sentence | Repair almost any single wrong tone | Nothing about your tones |
| A two-syllable word | Repair one tone from the other | Very little, and it will feel like a pass |
| One syllable, alone | Almost nothing | Roughly what you said, character for character |
| One syllable from a minimal pair | Nothing at all | Which of the two tones you actually produced |
How to run it in five minutes
Open any note-taking app, switch the keyboard to Chinese voice input, and work through a short list of single syllables that differ only in tone. Write down beforehand what you intend to say, because the whole value is in comparing the intention with the transcript, and doing that from memory afterwards is how people talk themselves into a better score than they got. Read the characters that come back rather than the pinyin, since the characters are the unambiguous record of which word it heard.
Two things make the result readable. Say each syllable on its own, with a clear pause on both sides, so the recogniser is not tempted to join it to its neighbour. And use syllables whose four tones are all real words, so that every mistake has somewhere to land: a wrong tone that is not a word gets snapped to the nearest one that is, which hides the error inside the same repair you were trying to escape.
What it still cannot do
This measures production, and production is the half of tone that most learners are not actually failing at. The harder half is perception: hearing which tone somebody else said, at speed, in a voice you have not heard before. A transcript of your own speech says nothing about that, and the two skills come apart in practice rather than moving together.
So this is a check, not a practice. It tells you where you stand on one half of the problem, on a handful of syllables, using a tool you already have, and it does that better than asking whether you were understood. Moving the other half takes the thing this cannot do: forced-choice listening with immediate feedback, on syllables you do not already know, which is where a known word supplies its own answer stops being a risk.
Frequently Asked Questions
What is the best way to check my Mandarin tones without a tutor?
Dictate isolated syllables into Chinese voice input and compare the characters that come back against a list you wrote down first. It is free, it takes five minutes, and removing the surrounding words removes the recogniser's ability to repair you. Tingyin covers the other half, which is hearing the tone rather than saying it, and no transcript of your own voice can measure that.
How do I stop speech recognition guessing what I meant?
Give it one syllable at a time with a real pause on either side, and pick syllables where all four tones are words in their own right. The guessing is the language model doing its job on context, so the only way to switch it off is to supply no context. A two-syllable word is already enough context for one tone to fix the other.
Is it safe to assume my tones are fine if Chinese dictation works?
No, and this is the most common way learners overrate themselves. Dictation working means the sentence was recoverable, which is a fact about Mandarin's redundancy and about the software, not about your pitch. The same sentence with two wrong tones very often transcribes perfectly.
What is the difference between speech recognition and a tone trainer?
Speech recognition judges what you produced; a tone trainer judges what you heard. They test opposite directions of the same skill, and progress in one does not automatically appear in the other. Tingyin plays a human recording and asks which of the four tones it was, which is a question dictation never asks you.
My phone writes the right characters, so why is my teacher correcting me?
Because your teacher is judging the sound and your phone is judging the sentence. Both are being honest; they are answering different questions, and the phone has more to work with. If the two disagree, the teacher is the one measuring the thing you are trying to improve.