{
 "run": "2026-09-10-tingyin-tone-measurement",
 "measured_on": "2026-09-10T19:05:37Z",
 "product": "Tingyin",
 "question": "Of the syllables in 624 publicly available recordings of Mandarin words spoken by 20 people, what share has a measured pitch contour whose shape matches the shape the five-level tone notation assigns to its tone?",
 "tier": "Standard",
 "cost_usd": 0,
 "wall_clock_min": 189,
 "sample": {
  "files": 624,
  "files_ok": 624,
  "nuclei_rows": 975,
  "nuclei_raw": 1212,
  "nuclei_merged": 975,
  "files_segmentation_agrees": 309
 },
 "files_on_disk": [
  "DATASET.json",
  "DATASET.csv",
  "LEDGER.jsonl",
  "PREREGISTRATION.md",
  "RUN-METADATA.json"
 ],
 "licence": "CC BY-SA 4.0",
 "licence_of_the_sources": {
  "CC BY-SA 4.0": 287,
  "CC0": 113,
  "Apache-2.0": 224
 },
 "instrument": {
  "environment": "packages/audio-gate in the Tingyin repository, under uv",
  "libraries": [
   "python 3.13.5",
   "numpy 2.4.6",
   "pyworld 0.3.5",
   "parselmouth 0.4.7",
   "librosa 0.11.0",
   "soundfile"
  ],
  "pitch_trackers": [
   "pyworld harvest+stonemask",
   "Praat autocorrelation via parselmouth",
   "librosa pYIN"
  ],
  "segmentation": "loudness-envelope nuclei, no claim about the syllable count"
 },
 "what_is_not_published": "The audio itself. The recordings belong to their speakers and to the two corpora; the dataset carries the SHA-256 of every prepared file, its speaker, its licence and the page it came from, so every measurement can be traced and no recording is redistributed.",
 "limits": [
  "Two volunteer corpora and twenty speakers; four speakers supply most of Lingua Libre and eight supply AISHELL-3, so nothing here is a sample of Mandarin speakers.",
  "Short isolated words and citation readings, not connected speech.",
  "Per-tone shares are computed only on the files where the nucleus count equals the claimed syllable count, because that index is what maps a nucleus to a tone.",
  "The syllable boundary is uncertain by the distance the two segmentation cues disagree by, and moving it moves a large share of the verdicts: see `replication`.",
  "The re-execution read one sentence of the written method the other way, about whether a syllable the segmenter cut in two is graded as one syllable or as the segmenter left it. That single switch moves the third tone share from 0.2069 to 0.0988, which is further than any interval here, so both readings are published and the switch is named: see amendment A-09 in the pre-registration. A run that grades its own dataset should decide which of the two readings it wants before quoting a third-tone figure.",
  "The corpus was fixed before the first measurement and that cannot be shown from the record alone: `corpus.json` and `tools.json` were written after the last row. Every row carries the SHA-256 of its input, so the corpus is recoverable from the dataset.",
  "The corpus was fixed before the first measurement and that cannot be shown from the record alone, because corpus.json and tools.json were written after the last row. Every row carries the SHA-256 of its input, so the corpus is recoverable from the dataset."
 ],
 "evidence": "LEDGER.jsonl holds every number the article states, each with the file and the path inside it that the number is read from; those files are published under evidence/, at the paths the ledger names."
}