Grew Up Hearing Mandarin, but Unsure of the Tones?
Growing up with Mandarin at home gives you a real head start on tones, and it lives in the words you already know. Take the word away and much of it goes: in a 2024 study, 29 heritage speakers learning new words told apart words that differed only in tone at chance, 50 percent, from the first block to the last, while words that differed by a consonant or a vowel rose to about 70.

The gap nobody expects you to have
If you grew up hearing Mandarin from your parents and then went to school in English, you probably understand far more than you speak, and you speak far more than you can explain. Then somebody asks which tone a word is, a teacher or a pinyin quiz or a cousin, and you cannot say. You know the word. You can say it correctly. You have no idea whether it is a second or a third.
It is tempting to take that as a sign that something is missing, or the opposite, that it does not matter because you clearly hear the language fine. Neither is quite right. The research on heritage speakers, people raised in a home where one language is spoken who then switched to another as their dominant one, shows an ear that is genuinely better than a classroom learner's and genuinely not a native speaker's, with the difference sitting in a predictable place.
Where the advantage actually lives
The clearest picture comes from Chang and Yao, in the Heritage Language Journal in 2016. They recorded Mandarin from 15 heritage speakers in the United States, born to Mandarin-speaking parents but speaking English most of the time, alongside native speakers and adult learners whose first language was English. Then 64 listeners from mainland China identified the tone of every recording, so the measure is how recognisable each group's tones are to a native ear.
The result split down the middle. On single syllables said one at a time, the heritage speakers' tones were less recognisable than the native speakers', and not generally more recognisable than the adult learners'. In words and phrases of more than one syllable they were easier to recognise than the learners' and no less recognisable than the native speakers'. The authors put the advantage down to connected speech: heritage speakers handled the way tones bend into one another inside a word the way native speakers do, which is exactly what you would absorb at a dinner table and never be taught.
That study measured how heritage speakers sound, not how they hear, and it is a small group. It still points at the shape that the listening studies then confirm: the skill is carried by familiar words in running speech, and it thins out when a syllable is taken out of that setting. The paper's own summary is careful about it. Early experience with the language can give an advantage on tone over adult learners, but it does not necessarily do so.
What happened when the words were new
Ge, Rato, Rebuschat and Monaghan, in Frontiers in Psychology in 2024, tested the listening side directly. Their 29 participants were born in Canada or the United States with at least one parent who spoke Mandarin natively. They heard invented two-syllable words, recorded by a native speaker, and learned what each one meant by seeing it paired with pictures across many trials, choosing one of two pictures each time. Some pairs of words differed a lot, some by one consonant, some by one vowel, and some only by the tone of the first syllable.
| The two words differed by | Accuracy by the last block (chance is 50%) |
|---|---|
| Several sounds | About 75% |
| One vowel | About 70% |
| One consonant | About 68% |
| Only the tone | About 53%, at chance throughout |

The tone contrast was not a subtle one. The first syllable carried either the first tone, high and level, or the fourth, a sharp fall: the two tones English-speaking learners generally identify most accurately, well ahead of the second and third. Heritage speakers who could plainly learn new Mandarin-sounding words could not use that difference to keep two of them apart. And how much Mandarin each person had used growing up did not predict how they did.
The authors' explanation is that English-dominant listeners weigh consonants and vowels more heavily when they are storing a new word, so that is what they attended to, and as a group they performed the way English speakers with no Mandarin do on the same task. Tone was available in the signal and the ear did not reach for it.
This is the heritage version of something every learner runs into: tone accuracy drops on words you do not know, because a known word hands you its tone before the audio finishes. A heritage speaker simply knows far more words, which hides the gap for longer and makes it more surprising when it shows.
The third tone, again
Where heritage listeners do trip on familiar material, it is mostly in one place. Chen and Shih, in the Journal of Chinese Language Teaching in 2021, played 42 native speakers, 21 heritage speakers and 25 adult learners two syllables separately and asked them to pick which of four recordings was the two said together the way a native speaker says them. The hard case is two third tones in a row, where the first is spoken as a second.
On that case 88 percent of native speakers chose the right answer, 71 percent of heritage speakers and 64 percent of learners. Heritage speakers were better than learners overall and closer to native, and the conditions they found hardest were exactly the third-tone ones, where the written tone and the spoken one stop agreeing and the textbook dip is the rarest sound the tone makes. Learners, by contrast, struggled with two-syllable sequences across the board.
Naming a tone is its own skill
Put the three studies together and the experience most heritage speakers describe makes sense. You say the word right because you learned the word, tone included, as one sound. What you never had to do was pull the tone out of the word and put a label on it, and nothing at home ever asked you to. School asks for it constantly, in dictation, in pinyin, in the tone marks a textbook prints over every syllable, and in looking up a word you heard but have never seen written.
That labelling step is learnable, and it is learned the same way as it is by anyone else: hear a syllable on its own, commit to one of four answers, find out at once whether you were right, and do it many times with many voices. The difference is that you are not starting from nothing. Your ear already treats tone as part of a word. The work is getting it to do so on syllables that arrive without a word attached.
Tingyin is a Mandarin tone-listening trainer that does only that: it plays a clip and asks which tone it was. Every clip is a human recording that carries its source, speaker and licence in a public manifest, 637 of them from eight speakers. Level 1 needs no account, level 2 is free once you sign in, and levels 3 and 4 are part of Premium. Because it tests hearing rather than speaking, it is aimed at precisely the half of the skill a heritage background leaves uneven. Hearing and saying tones are different skills goes into why the two come apart.
How to find out where you actually stand
Fluency in conversation is bad evidence here, because context repairs tone for listeners and speakers alike. A fairer check takes ten minutes. Listen to single syllables with no pinyin or characters on screen and name the tone of each before you see the answer. Then do the same with two-syllable words you do not know. If the first goes well and the second does not, you are looking at the pattern in the studies above, and the second test is the one to practise on. If the misses cluster on the third tone, that is the most common heritage pattern, and the third-tone page is the place to start.
Which of these Tingyin can test, counted from its own clips
We counted the 637 clips in Tingyin's audio manifest against the three findings above, because a trainer only helps a heritage speaker if it tests the part that is actually weak. It covers two of them and not the third.
| Level | What is in it | What it tests for a heritage speaker |
|---|---|---|
| 1 and 2 | 280 single syllables, 70 for each tone, all but one recorded by the same speaker | Naming a tone with no word attached, where the gap is |
| 3 | 237 everyday two-syllable words, such as 爸爸, 吃饭 and 电话 | Tone inside familiar words, where the head start already is |
| 3, a subset | 20 words written with two third tones, such as 可以 and 所以, answered as spoken: second then third | The third-tone case Chen and Shih found hardest |
The single-syllable levels are the useful part for most heritage speakers, and they are the free ones: level 1 needs no account and level 2 needs a sign-in. Level 3 is worth knowing about for a different reason. Its words are common ones, so a heritage speaker will usually know them, and a good score there confirms the advantage Chang and Yao measured. It does not test the new-word gap from the 2024 study, and nothing in Tingyin does, because every clip is a real word. The 20 third-tone words, all recorded by one speaker, are the one place the trainer reaches the hardest heritage case directly, and because they are marked as spoken rather than as written, answering third then third on one of them means you reached for the written tone instead of the one you heard.
Frequently asked questions
Do heritage speakers hear Mandarin tones like native speakers?
Better than adult classroom learners, not as reliably as native speakers. Chen and Shih (2021) found heritage speakers between the two groups on two-syllable tone identification, with their hardest cases on the third tone. The gap widens on words they do not know: in Ge and colleagues (2024), 29 heritage speakers could not use a first-tone versus fourth-tone difference to tell two new words apart, scoring at chance throughout.
Why can I say a word with the right tone but not tell you which tone it is?
Because you learned the word as a whole sound, tone included, and never had to separate the tone out and name it. Saying the word draws on memory of that sound. Naming the tone is a categorisation task on the syllable itself, and it is the task heritage speakers are least practised at, since nothing in conversation asks for it.
Does using more Mandarin at home make a heritage speaker better at tones?
Not necessarily, on the evidence so far. In the 2024 word-learning study, how much Mandarin each participant had used did not predict how well they told tone-only pairs apart. In Chang and Yao (2016), heritage speakers with more exposure were somewhat more recognisable in words of several syllables, but the two subgroups did not differ on single syllables. Both studies are small, so treat this as a pattern rather than a rule.
Which tone do heritage speakers find hardest?
The third. In Chen and Shih (2021), heritage speakers' weakest conditions were two third tones in a row and the shortened half-third that comes before other tones. Chang and Yao (2016) also found the third tone the hardest to recognise in every group's recordings, with native speakers' third tones easier to recognise than heritage speakers'.
I am a heritage speaker. Is an ear-training app too basic for me?
The early levels may feel easy, and that is useful information rather than wasted time: it tells you whether single syllables are already solid. Tingyin's level 1 needs no account, so you can find out in a few minutes. Where it earns its place is the single-syllable work in levels 1 and 2, syllables you cannot lean on a known word for, which is the part the research shows heritage experience does not supply on its own.