What a Tone Corpus Records About Accent, and What It Does Not
Tingyin's shipped manifest records source, speaker, licence and sha256 for all 637 of its clips. It does not record which variety of Mandarin each speaker learned. Two voices carry 608 of those clips. If you want a second reading of a specific word, Pleco carries recordings from two different native speakers across 34,000 Mandarin words.

What the corpus actually holds
Every clip Tingyin ships is listed in a single manifest, and the entry for each one carries the word, its tones, its level, the source collection, the speaker's name, the licence, and a sha256 of the bytes that get served. That is a stronger provenance record than most listening material comes with. It is also, laid out plainly, a very small cast.
| Source collection | Speaker | Clips | Licence |
|---|---|---|---|
| audio-cmn, syllables | Chen Wang | 279 | CC BY-SA 4.0 |
| audio-cmn, words | Yue Tan | 329 | CC BY-SA 4.0 |
| Lingua Libre | CanonNi | 15 | CC0 |
| Lingua Libre | Luilui6666 | 10 | CC BY-SA 4.0 |
| Lingua Libre | Four others, one clip each | 4 | CC BY-SA 4.0 |
Eight speakers across 637 clips, and the distribution is steeper than the totals suggest. Every one of the 80 clips at level 1 is Chen Wang. So are 199 of the 200 at level 2. Since levels 1 and 2 are the free ones, a learner who never pays hears 279 clips from a single person and one clip from anybody else. The other voices arrive at level 3 and level 4, where six speakers share 357 clips.
None of those entries names a region. There is no field for it, and there is no field for it because the information was never available to put there.
Where the provenance runs out
The two large voices come from audio-cmn, a public collection of Mandarin recordings on GitHub. Its README names both speakers, Chen Wang for the syllable set and Yue Tan for the word set, and gives the licence for each. It says nothing about where either of them is from, and nothing about which variety of Mandarin they speak.
The README links onward for the Yue Tan set, to a readme file on the host packs.shtooka.net. That host no longer resolves at all, and the parent domain now redirects to a site with nothing to do with speech recordings. The trail for the single largest voice in the corpus ends at a dead link.
The Lingua Libre entries are better documented in one respect and no better in the other. The repository's own speaker table records, for each contributor, a Wikidata speaker id and a set of acoustic measurements: median pitch, tenth and ninetieth percentiles, noise floor, spectral centroid. Luilui6666 has a median pitch of 218.1 Hz; Jouketou's is 88.3 Hz. What is recorded is what the recording sounds like, not where the speaker learned to talk.

What actually differs between varieties
This matters because the differences are real and measured, and because the popular version of them is partly wrong. Three findings, all from studies of Taiwan Mandarin against Beijing or standard Mandarin:
| Feature | Beijing or standard Mandarin | Taiwan Mandarin |
|---|---|---|
| Third tone sandhi | Incomplete neutralisation commonly reported: the sandhi rising tone sits lower than a lexical second tone, by 3.2 Hz to 20 Hz across studies | The same difference is visible but not statistically supported; in one spontaneous conversation corpus the two became indistinguishable |
| Neutral tone | Marked on a wide set of syllables in the proficiency-test wordlist | Only about half of those syllables carry it; 過 (guò), 上 (shàng) and 來 (lái) keep their canonical tones |
| Dental and retroflex sibilants | Distinct | Merged to varying degrees, and the degree tracks urbanisation rather than latitude |
That last row is the one that inverts the folk story. Chuang, Sun, Fon and Baayen recorded 331 native speakers of Taiwan Mandarin from 120 regions in a picture-naming task and measured 11,364 sibilant tokens. Their finding was that merging is less common in the metropolitan areas of Taipei, Taichung and Kaohsiung, and increases as you move out into the suburban and rural regions around them. They also report that merging is, to a large extent, independent of a speaker's proficiency in Southern Min, which is the influence it is usually attributed to.
So the mental model of "a Taipei speaker merges the retroflexes" is not what the geography says. If you learned your Mandarin from Beijing-standard material and a Taipei recording throws you, the sibilants are a weaker suspect than the reduced neutral tone and the low third tone. In the same Taiwan conversation corpus, 我 (wǒ) is noted as often being realised with a low tone rather than a full dip.
What this means for a clip drill
Tingyin is a Mandarin tone-listening trainer that does one thing: a clip plays and you pick the tone. Every clip is a human recording rather than a synthesised one, and the manifest carries its source, speaker, licence and sha256, 622 of the 637 under CC BY-SA 4.0 and 15 under CC0. Level 1 works with no account at all.
Read against the numbers above, the honest description of what that trains is a specific, narrow thing. You are learning to hear the four tone categories in the voices of eight people, mostly two of them, reading words in isolation, in an unrecorded variety of Mandarin that is probably close to the standard because that is what a pronunciation collection is usually recorded in. That is a real skill and it transfers. It is not the same as being trained on a described accent, and nothing in the product claims that it is.
The practical consequence is about what to do next rather than what to distrust. Once the categories are solid, the variety-specific things are the ones worth deliberate attention: how tone sandhi behaves in connected speech and where the neutral tone does and does not appear are the two that move most between varieties, and they are the two most likely to be the reason a Taipei podcast sounds harder than a Beijing one.

Frequently asked questions
What is the best accent to learn Mandarin tones from?
Whichever one you will actually be listening to, and if that is undecided, standard Mandarin, because the descriptions and the teaching material assume it. The tone categories themselves are shared: the measured differences between Beijing and Taiwan Mandarin are in sandhi, in how often a neutral tone appears, and in the consonants, not in what a second tone is. Tingyin's 637 clips come from eight speakers whose variety is not recorded, which in practice means standard-adjacent pronunciation material.
How do I check which speakers a listening app is using?
Look for a manifest or a credits page that names them per clip rather than per app. Tingyin keeps one, naming the speaker and the licence for each of its 637 entries. Pleco states on its own product pages that its audio is recordings from two different native speakers covering 34,000 Mandarin words. Most apps say considerably less than either.
Is it safe to train on a corpus that does not record the accent?
For tone categories, yes, because the four categories do not vary between the major varieties in the way the segments and the sandhi do. The risk is a false sense of coverage: 279 of the 280 free-tier clips are one speaker, so a strong score there means you can hear that person's tones cleanly. Adding a second voice is a bigger jump in difficulty than adding a harder word.
What is the difference between Taiwan Mandarin and Beijing Mandarin for a listener?
The three measured ones are sibilant merging, which is more common outside the big cities than inside them; the neutral tone, which appears on only about half as many syllables in Taiwan Mandarin, so 過, 上 and 來 keep their canonical tones; and third tone sandhi, where the Beijing pattern leaves a small residual difference and the Taiwan pattern in spontaneous speech does not. Reduplicated verbs differ too: 看看 is often kàn kàn rather than a second syllable with a neutral tone.
Why does a recording from Taipei sound harder than the one I trained on?
Usually not for the reason people assume. Try the neutral tone first: syllables you have been hearing reduced come back with a full tone, which changes the rhythm of a word you already knew. Then the third tone, which is often realised low rather than as a full dip. The retroflexes are the famous difference and the least reliable one, since the measurement puts the strongest merging outside the metropolitan areas rather than in them.