How Often Does the Mandarin Third Tone Dip?
Last updated: September 11, 2026
The first tone is level in 88.1% of the 118 graded syllables, the third tone dips in 20.7% of 58.
- Registered and confirmed: the dip landed inside 10% to 40%.
- Registered and unsettled: the fourth tone was expected to match best, at 70% or better. It clears that floor at 76% of 104, but which tone wins depends on how strictly level is drawn.
- The soft spot: moving every syllable boundary by the disagreement between the cues changes 32.6% of 1253 verdicts.
- Measured on 10 Sep 2026, over 624 public recordings.
How often each tone has the shape it is drawn with
The diagram every learner meets draws each tone as a shape: the first level, the second rising, the third falling and then rising again, the fourth falling. This measurement asks, syllable by syllable, whether the pitch actually does that. On the first tone it usually does, in 88.1% of the 118 syllables that reached a verdict. What counts as level is a decision rather than a fact, and a stricter test of it moves that figure a long way, which a later section shows.
The rest fall away from there. The fourth tone has its fall in 76% of 104, the second has its rise in 57.6% of 106, and the third tone has the fall-and-rise it is named for in 20.7% of 58. Each share has its own denominator, because each tone loses a different part of its syllables before it can be graded at all.
| Tone | Has that shape | Graded of claimed | Range by speaker | Range by item |
|---|---|---|---|---|
| First, level | 88.1% | 118 of 126 | 78.7% to 98.3% | 81.7% to 93.6% |
| Second, rising | 57.6% | 106 of 114 | 37.9% to 68.2% | 47.1% to 69% |
| Third, dip and rise | 20.7% | 58 of 81 | 0% to 30.2% | 10.4% to 32.7% |
| Fourth, falling | 76% | 104 of 129 | 62.5% to 83.7% | 67% to 84.2% |
The two ranges say what the same measurement would have produced on a different draw of the same size, resampled once by speaker and once by item. They are wide. On the third tone the range by speaker reaches down to not one graded syllable, which is what a share resting on 58 nuclei from a handful of voices looks like when the voices are drawn again, and the range by item on that row is 10.4% to 32.7%.
Every share here counts only the syllables that reached a verdict, and the tones do not lose the same amount getting there. The third tone loses the largest part: 28.4% of its claimed syllables were never graded, against 6.3% on the first tone, 7% on the second and 19.4% on the fourth. The syllable hardest to cut cleanly is the long dipping one, and the splitting counts further down show it.
Interpretation
A discard rate that differs by tone is not neutral. The syllables that survive to be graded are the ones that were easy to cut, and if the hard ones are disproportionately third tones, then the third tone's share is measured on its easier cases. That pushes the dip share up rather than down, which makes the figure below more likely generous than harsh.
The third tone against the range that was registered
Before any audio was measured, this run wrote down what it expected of the third tone: fewer than half of its nuclei would carry the fall-and-rise, and the share would land between 10% to 40%. The measured share is 20.7%, on 58 graded nuclei. That is inside the registered range, near the bottom of it.
20.7%
of 58 graded third-tone syllables carry the fall-and-rise the diagram draws for them. It is the shape the tone is taught by, and in these recordings it is the exception.
That figure carries an instrument problem, and it belongs beside the number. The rise at the end of a third tone often happens in creaky voice, where a pitch tracker stops returning a pitch, so the part of the syllable the notation describes is the part most likely to be missing from the window. The separate measurement of this corpus found the third tone reading as a plain fall, a shape the notation does not write, in 74% of its 192 third-tone syllables. Neither design can separate a real low-falling allotone from a truncated rise, and the other method went at the question directly:
- Split those same third-tone syllables into quarters by how much of each one a tracker could measure, and the dip rate rises with the coverage: 2.1% to 12.5% to 20.8% to 18.8% across the quarters. It climbs steeply out of the least-covered quarter and has not settled by the best-covered one, which still leaves part of the syllable outside the window. The dip share above is a floor and not an estimate.
- The first tone, split the same way, sits at 54.5% to 53% across the same quarters. It shows no gradient, and with almost nothing in its least-covered quarter it could not have shown one.
- The fourth tone runs the other way: 90% to 69.5% from its least-covered quarter to its best-covered one. A longer window on a falling tone starts to take in the flat bottom of the fall, which costs it matches rather than winning any. A gradient on its own is therefore not evidence about the notation, which is why both controls are here.
The rival explanation lives in where the syllable sits. If the dip is a shape that only comes out when a syllable has room, it should be commonest on one read on its own, and it is: the rate runs 15.7% to 11.7% to 20.8% across the positions, with the highest of them in isolation. A window that cuts the rise off produces that same order, so this separates neither explanation from the other.
That same measurement ran a check this design never ran, on the 1085 syllables it graded. A classifier trained on every speaker but one, and tested on the speaker it had not heard, recovers the tone from the measured contours 63.9% of the time, against 29.4% for naming the commonest tone every time. Replace what it learned with the shapes the notation draws and it falls to 53.7%. On the third tone, keeping only the height and discarding all the movement still recognises it 76% of the time.
Where the recording is cut cleanly, the dip turns up more often. On files whose segmentation agrees with the claimed syllable count it is 23.5% of 81 third-tone items, against 10.7% of 140 where the two disagree. Of 221 third-tone items, 35 items were cut into more units than the word claims, and on those the trackers came back with a rise on 35.1% of 165 verdicts. A dip cut through the middle is a fall and then a rise, and the rise on its own reads as a second tone.
Further readings of this same tone exist, and they belong here rather than in a footnote:
- Running the whole measurement again from the start returned 10.2% on 49 graded nuclei.
- Moving every syllable boundary by the distance the cues disagree by returned 17.6% on 51.
- The separate measurement of the same corpus returned 13.5% on 192.
The direction holds in every one of them: the dip is the rarest of the drawn shapes. The share does not hold to the precision this design claims: stated plainly, the third tone did not replicate to within the design's own precision, while its direction did, and every replication is printed beside it in the table above. So 20.7% is the middle of a spread rather than a figure to quote on its own.
The expectation about the order, and why it is still open
The other registered expectation named the fourth tone as the best matched of the tones, at 70% of its nuclei or better, with the first tone behind it and the second tone last of all. The floor holds: the fourth tone has its fall in 76% of 104. On this run's reading the order does not. The first tone is higher, at 88.1%, and the tone that comes last is the third, not the second.
The separate measurement of the same corpus puts them back the other way round. It reads the fourth tone at 78.1% of 319 graded syllables and the first at 57.9% of 292, which is the order that was registered. The difference is not a mistake in either: the two draw the level shape differently, and a strict test of flatness costs the first tone a great deal and the fourth tone almost nothing.
57.9%
of 292 first-tone syllables are level on the stricter of the two tests, against 88.1% of 118 on this one. Which tone matches its notation shape best is a question about where that line is drawn, and this run does not settle it.
There is a further reason to leave it open. The registration and the summary approved alongside it name different tones here, and the registration governs: it named the fourth. Read that way, this run's reading is right in its floor and wrong in its order. Read the other way, with the first tone named, it holds, since 88.1% clears the same floor of 70%. Both readings are published, because choosing the flattering one after seeing the answer is how a registered expectation stops meaning anything.
Interpretation
What is left is not a tidy refutation, and saying so is the point. The fourth tone was expected to win because it is short, loud and steeply falling, which is what makes it easy for a person to hear. On a loose reading of level the first tone wins instead, because a flat stretch of pitch is the easiest thing for a tracker to agree on, and on a strict one it loses badly. A ranking of shapes measures the method as much as it measures the language, and the one place both methods agree is that the dip comes last.
What moved the answer
Every share above rests on a decision about where each syllable starts and stops, and the cues that make that decision do not agree with one another. On the volunteer recordings they sit a median of 75 milliseconds apart, over 659 nuclei, and on the read utterances 160 milliseconds apart, over 553. Moving each boundary by that distance changes the shape verdict on 32.6% of 1253 nuclei, and changes which shape wins on 32.3% of them.
32.6%
of 1253 shape verdicts change when the syllable boundaries move by the distance the segmentation cues disagree by. Every share in this article sits on top of that one decision.
The same disagreement shows up in the counting. Set what the segmenter found against what each word claims, and the surplus and the shortfall both appear:
- First tone: the segmenter found 16 nuclei beyond the 324 syllables it is claimed on, and came up 83 nuclei short.
- Second tone: 28 nuclei beyond the 319 claimed, and 98 nuclei short.
- Third tone: 37 nuclei beyond the 234 claimed, and 81 nuclei short. That is the smallest claimed base of the tones and the largest surplus of them.
- Fourth tone: 13 nuclei beyond the 369 claimed, and 115 nuclei short.
- These are directions of error rather than rates: the count is taken per file and charged to every tone that file contains, so the tones overlap and the totals do not add up.
A separate session was handed the question and the recordings and nothing else: not this method, not this dataset, not this answer. It wrote its own method down first, measured the whole corpus its own way, and found 54.5% of 1085 syllables matching their notation shape. Its test of a match is not this one, so its figures are comparable with these in their order and not in their level.
The difference in level is a difference in design rather than in the language. This run graded 386 nuclei in the same recordings: it finds the syllable boundaries itself and keeps only the files where the count it finds matches the count the word claims. The other design places the boundaries against the count it already knows, and grades what it places.
| Tone | This run | Run again from the start | Boundaries moved | A separate measurement | Over-split left unjoined |
|---|---|---|---|---|---|
| First, level | 88.1% of 118 | 90.9% of 110 | 74.8% of 107 | 57.9% of 292 | 84.6% of 195 |
| Second, rising | 57.6% of 106 | 48.4% of 97 | 46.6% of 88 | 52.1% of 282 | 43.8% of 169 |
| Third, dip and rise | 20.7% of 58 | 10.2% of 49 | 17.6% of 51 | 13.5% of 192 | 9.9% of 81 |
| Fourth, falling | 76% of 104 | 75.3% of 97 | 75% of 100 | 78.1% of 319 | 76.6% of 167 |
The second column is the same method run again from the start, the third is that same run with every syllable edge moved, the fourth is the separate measurement, and the fifth throws one switch in this run's own pipeline. The fourth tone holds across all of them. The first does not: 88.1% here and 57.9% under the stricter test of level.
The third tone is lowest everywhere and steady nowhere: 20.7% here, 10.2% on the re-run, 17.6% with the boundaries moved, 9.9% with the over-split syllables left unjoined, and 13.5% on the other method. Its direction is stable and its share is not, which is why every reading is printed instead of the friendliest of them.
110 recordings
of the 624 recordings are joined although the segmenter had already cut them correctly, which is the largest thing this measurement decides for itself rather than reads off the audio.
The fifth column is the one worth understanding, because it is the widest spread in the table and it is not a measurement at all. The segmenter cuts syllables by where the energy falls, and on a word of one syllable it sometimes finds two nuclei. The pipeline joins the second nucleus back into the first before grading, because the word is one syllable and the merge recovers it. Leave them apart instead and the third tone falls from 20.7% of 58 to 9.9% of 81, while the fourth tone barely moves.
The merge is also the widest thing this measurement does on its own authority. The function it comes from is a test of whether an over-count can be explained, and this run applies it wherever it points at a syllable, not only where the segmenter found too many. Applied that way it takes the number of recordings whose measured syllable count matches the claimed one from 399 recordings to 309 recordings, and on the second of the two corpora from 108 recordings to 34 recordings. Those 110 recordings were cut correctly before the join and are joined anyway, so anybody reusing this data should decide the rule for themselves first.
One more check leaves the answer where it was. Dropping every tracker that found a pitch in less than half the frames of a nucleus leaves the first tone at 87.8% of 115 and the third at 19.4% of 36, against 88.1% and 20.7% here. The shares barely move, and the denominators fall.
Interpretation
Which of the two readings of the over-split syllable is right is not a question about the recordings. It is a question about what the sentence describing the method means, and it moves the headline further than any interval printed here. What survives all of this is one ordering and a gap, not an exact share: the dip is the shape a tracker finds least often in every version of the measurement, and which shape it finds most often is a matter of where the test is set. Anybody quoting a single number for the dip should carry the spread with it.
Does the speaker move the pitch more than the tone does?
This was going to be the other side of the question, and it was withdrawn before anything was measured. What follows carries no registered direction: it was not predicted, and must not be read as though it had been.
Inside a single recording, the median distance between two different tones spoken by one person is 3.1 semitones, over 50 recordings. Across recordings of the same word by different speakers, the median distance is 7.6 semitones, over 192 pairs. The ledger records those as different statistics over different units, so they are printed beside each other here and not subtracted.
The separate measurement reaches this axis by a route that never compares two distances. Run its classifier on the pitch as recorded, with each speaker's own range left in, and the tone stops being recoverable at all: 29.5% of 1085, which is the baseline a classifier reaches by naming the commonest tone every time.
- The speakers who clear the recording threshold sit 14.7 semitones apart end to end, over 5 speakers.
- Between the quartiles the same speakers sit 3.1 semitones apart, which is the spread that the outermost voices are not driving.
- Every contour here is measured inside its own recording rather than against a fixed scale, because a register difference of that size would otherwise swamp the shape.
Interpretation
The larger of the two medians is the one between speakers, and it does not settle the question, because the two are not measured over the same units and the ledger says so. What can be said is that both distances are big enough to matter to a listener, and that this run was never built to arbitrate between them once the expectation was withdrawn.
What the corpus is, and what it is not
The recordings are public and none of them was made for this measurement: 400 recordings of volunteer words and 224 recordings of read utterances from an open speech corpus, 624 recordings in all, from 20 speakers. Between them the words claim 1303 syllables.
The segmentation agrees with the claimed syllable count on 309 files, which is 49.5% of 624, and that is the first thing to know about every share above.
The per-tone shares are counted only on the files where the two agree, because a share counted over syllables the segmenter never found is a share of something else. That restriction is what leaves the third tone with the thinnest base in the article, 20.7% on 58 graded nuclei, and it is why the ranges on that row are the widest.
- The neutral tone is in the recordings and not in this measurement, because the notation draws no shape for it to match. The words claim 57 syllables of the 1303 in the corpus. The segmenter came up 25 nuclei short of that base and found none beyond it. Of that base, 24 syllables carried enough voiced pitch to measure at all, and even those are described rather than graded: the notation gives the neutral tone no shape, so a contour on one of them cannot match anything and cannot fail to.
- These are isolated words and citation readings rather than conversation. Nothing here describes what a tone does in a sentence spoken at speed.
- The speakers are volunteers and read-speech contributors, not a sample of Mandarin speakers. The dataset names every one of them.
- Two recordings in the read-speech corpus pair a word whose transcript does not carry the tone that was sung, and the pipeline cannot detect either. The count is small and its direction is unknown, so the run publishes it as a limitation rather than repairing a third party's transcript.
Tone is recorded as it is spoken rather than as the dictionary writes it. Of 1303 measured syllables, the spoken tone differs from the dictionary tone on 33 syllables. The one systematic way that label could be wrong is third-tone sandhi, where a third tone before another is said as a rise instead. The context does occur. In the 400 recordings that carry a dictionary reading, 17 syllable pairs syllable pairs read as a third tone followed by another, and in every one of them the spoken label already records the rise, so no sandhi-affected syllable is graded as a third tone.
The read-speech corpus carries no dictionary reading at all, so the question was not examined there, which is a gap rather than a clean result.
The data
Everything this run measured was measured on 10 Sep 2026 and comes out of one published file, and the readings taken from the separate measurement are named as such wherever they appear. A reader who doubts a figure here should open the file rather than take the sentence around it on trust.
- A row for each nucleus, carrying every tracker's own verdict and the majority of them that the shape rests on.
- A row for each recording, carrying its source, its speaker, its licence and the page it was published on.
- The limits of the measurement, written out in the file itself.
- The measurements, one row per nucleus, as JSON.
- The same table as CSV, for a spreadsheet.
- What was written down before the first file was opened, with every later change to it appended and dated.
- The run record: the instruments and their versions, the counts per group, and the limits.
- Every number this article states, each with the file and the path inside it that it was read from, and the files themselves under
evidence/. - The dataset page, which lists all of the above in one place and is the address to cite.
Everything above is published under CC BY-SA 4.0. The measurements are derived from Lingua Libre recordings, themselves CC BY-SA 4.0, and from AISHELL-3, under the Apache License 2.0. Every contributing recording is named individually in the credits, with its speaker, its licence and the page it was published on. The preparation applied to each recording, trimming, peak normalisation and an appended silence, is a modification and is declared as one. Measured on 10 Sep 2026.
Tingyin plays human recordings and asks which tone you heard. Why the textbook dip is the third tone's rarest sound is the same argument written for a learner rather than for a reader checking the numbers, and it now has 20.7% behind it.
Frequently asked questions
Does the Mandarin third tone really dip?
In these recordings it usually does not. The fall-and-rise the diagram draws turns up in 20.7% of the 58 graded third-tone syllables. Where the recording is cut cleanly it is commoner, 23.5% of 81 third-tone items, against 10.7% of 140 where the cut and the claimed syllable count disagree.
Which Mandarin tone matches its textbook shape most reliably?
It depends on how strictly the level shape is drawn, and the two measurements of this corpus disagree. This run puts the first tone top, at 88.1% of 118 graded syllables against 76% of 104 for the fourth. The separate measurement, whose test of level is stricter, puts the fourth top at 78.1% of 319 and the first at 57.9% of 292. The expectation registered before the run named the fourth, at 70% or better. Both readings agree the third tone is last.
How much can these shares be trusted?
Less than a single figure suggests. Moving the syllable boundaries by the distance the segmentation cues disagree by changes 32.6% of the 1253 shape verdicts. Leaving the syllables the segmenter split in two unjoined, instead of grading them as one syllable as this run does, moves the third tone from 20.7% of 58 to 9.9% of 81, a wider spread than any interval here. The direction is steady across every version, so read the ordering of the tones as the result and treat any one share as the middle of a spread.