```
run_id:        2026-09-10-tingyin-tone-measurement
product:       tingyin
tier:          standard
written_at:    2026-09-10T14:05:00Z
```

## 1. Question

Of the syllables in 624 publicly available recordings of Mandarin words spoken by 20 people, what
share has a measured pitch contour whose shape matches the shape the five-level tone notation
assigns to its tone — and is the difference between two speakers saying the same word larger than
the difference between two different tones spoken by one person?

## 2. Method

**Inputs and where they come from.** Two public corpora, both already described by committed
manifests in the Tingyin repository, both downloaded on demand and never committed:

- **Lingua Libre** — 400 files, 289 repertoire items, 468 claimed syllables, 12 speakers, from
  `upload.wikimedia.org`. The list is `fixtures/lingua-libre/manifest.json`; every file carries its
  licence, its speaker and its Commons page.
- **AISHELL-3** — 224 single-word utterances, 8 speakers, over the spans in
  `scripts/audio/corpora/aishell3.json`, from the Hugging Face mirror `AISHELL/AISHELL-3`.

Fetch with `uv run --project packages/audio-gate python scripts/audio/corpora/fetch.py`, which
trims, peak-normalises to −3 dBFS, appends 200 ms of silence and writes 16-bit PCM. Preparations
are a modification of the source and are declared as one.

**Tool or command per step.** Nothing is installed and nothing is paid for. Every step runs inside
the existing environment `packages/audio-gate` (`uv`), which already carries pyworld, parselmouth
and librosa; the module `tingyin_gate` supplies the primitives.

1. `metrics.load_audio` reads each file to mono float64.
2. `metrics.frame_rms_db` computes the RMS envelope on the 5 ms grid; the grid length is `n`.
3. `metrics.trackers_of(path, x, sr, n, ['pyworld','praat','pyin'])` produces three independent F0
   tracks on that one grid. They are computed because no single algorithm's failure may decide a
   verdict alone.
4. `segment.segment_energy` finds the nuclei from the RMS envelope alone, blind to the claimed
   syllable count and to the claimed tone.
5. For each nucleus, `segment.contour_window(a, b, f0s, rms)` picks the span that carries the tone.
6. `metrics.features(f0, wa, wb)` returns, per tracker, the contour resampled to 20 points and
   time-normalised, plus `onset_st`, `offset_st`, `net_st`, `range_st`, `t_min`, `t_max`,
   `fall_leg_st`, `rise_leg_st`, `up_leg_st`, `down_leg_st`, `body_*`, `median_hz`, `dur_ms`.
7. `metrics.shape_of` and `tones.satisfied_tones` turn those into a shape verdict and the set of
   tones the contour is consistent with.
8. A per-nucleus verdict uses the gate's own 2-of-3 combination across trackers.

**Two measurement axes, on two different scales, and this is deliberate.**

- **Axis 1, the shape, is measured on the syllable-normalised contour.** `features` reports every
  value in semitones around that syllable's own median pitch, so a contour is compared with the
  notation's shape and never with its absolute pitch. The five-level notation is read as a shape:
  tone 1 level, tone 2 rising to a late high, tone 3 falling then rising with its minimum inside
  the syllable, tone 4 falling from an early high. The comparison uses sign and size of the net
  change, `range_st`, and the positions `t_min` and `t_max` — never an absolute level, because the
  mapping from the notation's five levels onto semitones is not defined by anything and inventing it
  would put an arbitrary constant inside the result.
- **Axis 2, the speakers, is measured in hertz and never normalised.** Axis 1 divides each
  syllable by its own median pitch, which by construction deletes the speaker. Axis 2 therefore
  compares raw `median_hz`, raw `range_st` and raw duration across speakers, and the same three
  quantities across tones within one speaker, on the 104 Lingua Libre items that carry two or more
  speakers.

**Repetitions per input.** Three F0 trackers per nucleus, which is the repetition inside the unit;
one pass over each file; and a separate re-execution session repeating the whole pipeline on the
same input list to write its own dataset.

**Output file and its shape.** `data/units.jsonl`, one row per nucleus, appended the moment it
exists, never reconstructed. Fields: `run_id`, `unit_id` (`<file>|<nucleus index>`), `item_id`,
`hanzi`, `pinyin`, `claimed_tones`, `syllable_position`, `corpus`, `speaker`, `license`,
`file_sha256`, `nuclei_in_file`, `claimed_syllables`, `segmentation_agrees`, `nucleus_a`,
`nucleus_b`, and per tracker `{shape, net_st, range_st, t_min, t_max, dur_ms, median_hz,
n_voiced}`, plus `verdict_tone` and `verdict_kind`. `data/files.jsonl` carries one row per file with
the fetch status, the RIFF check and the byte count.

**What is discarded, and on what rule.** A nucleus is discarded, and the fact is counted, when
(a) fewer than two of the three trackers return a contour on it, (b) its contour window is shorter
than `MIN_SYL_MS`, or (c) `features` returns `None` for every tracker. Discarded rows stay in
`data/discarded/` with their reason. **Nothing else is discarded, and no nucleus is ever dropped for
disagreeing with its claimed tone.**

**The denominator.** It is the number of nuclei that reached a verdict, not the number of files and
not the number of claimed syllables. The share of claimed syllables that reached a verdict is
published in the article beside every per-tone figure, because the segmentation step is the one that
can silently remove exactly the syllables that would change the answer.

## 3. Expectation

Three, each with a range, each grounded where the repository already carried the number.

1. **Tone 3 rarely shows its textbook dip.** Fewer than half of the tone 3 nuclei measured here will
   contain a fall followed by a rise inside the syllable, and I expect the share to land between
   **0.10 and 0.40**. This is the direction Tingyin's own published article claims without a number.
2. **Tone 4 is the most reliably shaped tone**, matching the notation's fall in **at least 0.70** of
   its nuclei, and tone 1 the second most (a level contour is the easiest thing to measure), while
   tone 2 is the least reliable of the four. Tone 2 and tone 3 are the pair the literature and our
   own gate both single out, and the gate's own recall is weakest there.
3. **The speaker moves the pitch at least as much as the tone does.** Measured on the manifest's own
   per-speaker statistics, the 12 Lingua Libre speakers' median pitches span **17.29 semitones**
   end to end and **8.23 semitones between the quartiles**, while the pitch movement a tone carries
   inside one speaker is a handful of semitones. So I expect the between-speaker difference on a
   shared word to **exceed** the between-tone difference within a speaker, on **more than half** of
   the 104 shared items.

## 4. Falsifier

- Expectation 1 is wrong if **more than half** of the tone 3 nuclei measured contain the dip.
- Expectation 2 is wrong if some other tone matches its notation shape more often than tone 4, or if
  tone 4 falls below 0.70.
- Expectation 3 is wrong if the between-speaker difference exceeds the between-tone difference on
  **fewer than half** of the shared items, in which case the article is about the opposite of what it
  predicted and says so.

A falsified expectation is a result and is rendered in the article in its own section, quoting the
range that was registered here.

## 5. Abandonment condition

- If segmentation agrees with the claimed syllable count on **fewer than 30 per cent of files**, the
  run stops and records `NOT_MEASURABLE`, saying which files it lost and what would have been needed.
- If the Contour window is shorter than `MIN_SYL_MS` on more than half of nuclei, the run stops and
  records `METHOD_BROKEN`, saying which stage broke.
- If neither corpus can be downloaded after three paced attempts with the RIFF check, the run stops
  and records `NOT_MEASURABLE`, not `FAILED`, because the question remains answerable elsewhere.

## 6. Estimate

- **What the run needs from Jakub:** one decision. The dataset is derived from 287 files under
  CC BY-SA 4.0, so it must be published under **CC BY-SA 4.0** and not under the CC0 that
  `publishing.md` records from the Viallo run (`NEEDS_USER`). Everything else it needs is nothing:
  no account, no payment, no device, no cooperation from anyone.
- Machine time: roughly 2 hours of measurement over 624 files, plus about 15 minutes of paced
  downloading.
- Money: **0 USD.** No paid interface is called.
- Claims expected in the ledger: about 12 measured, 3 to 5 sourced, and no own-data claims.

## 7. Adversary dispositions

Twenty-one objections, all raised before any measurement of this run, all read against the
repository's own committed files. Correction number in brackets points at amendment A-02.

**Section 1 — inputs that would make the conclusion false**

- **1.1 `claimed_tones` names two fields that disagree on 24 syllables — changed.** `claimed_tones`
  is `tones`, the post-sandhi field, and the 24 syllables where it differs from `citationTones` are
  published as their own line with both readings. [9]
- **1.2 AISHELL-3 is not what the sample statement described — changed.** Verified independently:
  224 rows, 70 disyllables, 73 trisyllables, 81 quadrisyllables, no monosyllables, 683 claimed
  syllables, content that is music-application search phrases and singer names. The note inside
  `aishell3.json` asserts the opposite of what the repository's own reference file records. The
  inputs are re-described and the label file is named. [8]
- **1.3 Two named AISHELL-3 rows carry a tone that does not match the word — accepted as a
  limitation.** I checked both: `SSB03150254` labels 周界伦 where the singer is 周杰伦, and
  `SSB03150038` pairs the surname reading of 单 with the common word's tone. The pipeline cannot
  detect either. The count is small and its direction is unknown, so the article states the
  limitation in its own sentence rather than the run repairing a third party's transcript.
- **1.4 One file, two claimed tones, one `unit_id` — changed.** The loop deduplicates on
  `localFile` exactly as the download loop already does, and `unit_id` carries the claiming item. [3]
- **1.5 The neutral tone cannot be graded and has no discard category — changed.** Tone 5 leaves
  the denominator before the run and its count is published. [2]
- **1.6 Steps 3 and 4 are in the wrong order — changed.** Confirmed by reading `gate.py:357-360`,
  which segments first and comments why. [1]
- **1.7 `F0_FLOOR` is global and creak cannot be told from an octave error — changed.** The floor is
  set per speaker from the manifest's own `f0P10Hz`, and the frames each nucleus loses to `despike`
  become a published column instead of a silent one.

**Section 2 — confounds**

- **2.1 My own amendment A-01 drew the wrong consequence — changed, and this is the most valuable
  objection in the report.** A-01 said the dipping syllable is split and therefore discarded.
  Verified independently: nothing is discarded. On the 245 Lingua Libre monosyllables the segmenter
  over-splits 33 of 63 tone 3 files against 10 of 68 tone 4, and the 63 tone 3 syllables yield 97
  nuclei, all of which reach a verdict. So A-01's own safeguard would have reported **154 per cent
  coverage for tone 3 against 115 per cent for tone 4**, making the worst-handled tone look the best
  covered. The guard points away from the defect it was written for. Replaced by a signed count
  plus `segment.artefact_nuclei`. [5]
- **2.2 AISHELL-3 loses syllables rather than splitting them — changed.** 113 of its 224 files yield
  fewer nuclei than syllables. For every such file the merged syllables are resolved from the pinyin
  and an interval is published whose ends assume all, and none, of them would have matched.
- **2.3 Onset voicing is unbalanced across tones by 26 points — changed.** Every per-tone share is
  split three ways by onset class. [6]
- **2.4 A tracker voicing a third of the window casts a full vote — changed.** Coverage is published
  per tone per tracker and the vote is re-run with a floor. [7]
- **2.5 Speaker identity is confounded with the recording chain — changed.** Axis 2 is computed
  twice, once on `median_hz` and once on `range_st`, and the article says that a difference
  appearing only in `range_st` is the chain, not the speaker.
- **2.6 The between-tone term varies the rime as well as the tone — changed.** It is computed paired
  inside one recording, on the 70 disyllables carrying two distinct non-neutral tones.

**Section 3 — what the author was not thinking of**

- **3.1 Both halves of the comparison are the same statistic — changed, and it forces a
  re-registration.** See the amendment below the dispositions.
- **3.2 The registered 17,29-semitone span rests on two single-file speakers — changed.** Verified:
  the endpoints report an interdecile F0 spread of 0.00 semitones, which no speaker producing four
  tones can have. The speaker range is computed only over speakers with at least 10 usable
  recordings, and every per-speaker figure carries its count. [12]
- **3.3 No uncertainty, and the units are not independent — changed.** Bootstrap by speaker and by
  item, printed beside every registered range. [11]
- **3.4 Normalising by the syllable median deletes the notation's levels — changed.** A second
  column is added in which each syllable's onset and offset are expressed in semitones relative to
  that speaker's own median, which tests the notation's ordinal predictions without inventing a
  mapping from its five levels to semitones.
- **3.5 `shape_of` and `satisfied_tones` differ on tone 3 by design — changed.** The numerator is
  declared, and the other verdict is published beside it. [4]
- **3.6 Five gate outcomes, two in the question, mapping undeclared — changed.** [10]
- **3.7 The re-execution replication measures the disk — changed.** The objection is correct and the
  replication is redesigned to perturb the boundaries.
- **3.8 The second half of the question is not answerable as specified — changed.** The estimator is
  replaced as the adversary specifies. The expectation attached to it is withdrawn rather than
  re-set; the reason is in the amendment.

## 8. Amendments

Append-only. None yet.

### Amendment A-01 — the discarded syllables are not discarded at random

`written_at: 2026-09-10T14:45:00Z` · `post_hoc: false` · `author: the run`

**What was seen before this amendment.** Not a result of this run. The gate's own committed
reference measurements, `scripts/audio/corpora/reference/seg_native_before.json`, 644 rows from the
same two corpora, were read to size the sample. They carry, per item, the claimed syllable count and
the nucleus count `segment_energy` found. Recomputed here:

| group | agrees with the claim |
|---|---|
| all 644 items | 414 (64.3 %) |
| monosyllables | 176 of 245 (71.8 %) |
| two or more syllables | 238 of 399 (59.6 %) |
| items containing tone 1 | 174 of 263 (66.2 %) |
| items containing tone 2 | 160 of 272 (58.8 %) |
| **items containing tone 3** | **119 of 228 (52.2 %)** |
| items containing tone 4 | 175 of 287 (61.0 %) |
| items containing the neutral tone | 35 of 56 (62.5 %) |

**The change.** Two things are added, both before any measurement.

1. **Every per-tone figure is published with its own discard rate.** The article reports, for each
   tone, the number of nuclei that reached a verdict out of the number of syllables claimed for that
   tone, in the article body and not only in the dataset.
2. **A discriminating test is added.** For tone 3 items the run computes the share containing a fall
   followed by a rise separately in the files where segmentation agreed and in the files where it did
   not. If the dips concentrate in the discarded half, the registered expectation 1 is not a finding
   about Mandarin and the run says so in place of the finding.

**The reason, and it is a mechanism rather than a correlation.** A tone 3 carrying its full dip is
longer and has two pitch events, so its loudness envelope has two peaks where a level or falling
tone has one. The segmenter splits nuclei at dips in the envelope with sufficient prominence, so
**the syllable that contains the dip is the syllable most likely to be split in two and therefore
discarded.** The bias runs in the same direction as the expectation: it removes exactly the tone 3
syllables that would falsify "the dip is rare". Left uncorrected, expectation 1 could be confirmed
entirely by the instrument losing the evidence against it.

**What this does not change.** The question, the inputs, the tools and the estimators stand. The
expectation ranges stand as registered. What changes is that the run must show its work on the
denominator per tone, and must test the direction of the loss rather than only its size.

### Amendment A-02 — what the adversary changed, and the re-registration it forces

`written_at: 2026-09-10T15:40:00Z` · `post_hoc: false` · `author: the run, on ADVERSARY.md`

Twenty-one objections, raised before any measurement of this run. Dispositions are in §7. The
corrections that change the method itself:

1. **Steps 3 and 4 swap.** The gate's own `gate.py` segments before tracking and says why; the
   method as registered did the opposite and passed no spans, so every multi-syllable file got a
   file-global octave reference. `segment_energy` now runs first, its span set goes to
   `trackers_of`, and `contour_window` is computed from the syllable-local reference.
2. **Tone 5 leaves the denominator before the run.** 36 Lingua Libre and 17 AISHELL-3 syllables are
   neutral. `satisfied_tones` iterates `(1, 2, 3, 4)` only, so a neutral syllable can never be
   satisfied and every one of them is a mismatch by construction. They are counted and published,
   and they are not in any per-tone share.
3. **The measurement loop deduplicates on `localFile`.** 30 of the 400 files are claimed by more
   than one item, one of them with two mutually exclusive tones. `unit_id` becomes
   `<file>|<nucleus index>|<claiming item>` and the file's audio is measured once.
4. **The tone-3 numerator is declared.** `shape_of == DIP-RISE`, because that is the notation's
   214. `satisfied_tones` is published beside it as a separate column, never in place of it, because
   the gate's own comment records that requiring the rise "costs 23 of the 27 claimed tone 3s".
5. **The per-tone coverage number is signed, not a ratio.** A nucleus count over a claimed-syllable
   count reports 154 per cent for tone 3 against 115 per cent for tone 4, because the segmenter
   splits the dip rather than discarding it. Over-split and under-count are published separately,
   and `segment.artefact_nuclei` is applied to the over-split nuclei before grading.
6. **Every per-tone share is additionally split by onset class** — voiceless obstruent, sonorant,
   zero — because the vocabulary's onsets are unbalanced across tones by 26 points.
7. **Tracker coverage is published per tone per tracker** as `n_voiced / (dur_ms / 5)`, and the
   2-of-3 vote is re-run with a coverage floor to show whether the shares move.
8. **AISHELL-3 is described correctly, and its labels are named.** It is 224 utterances of 2 to 4
   syllables, 683 claimed syllables, no monosyllables, content that is mostly music-application
   search phrases and singer names. Its hanzi, pinyin and tones come from
   `scripts/audio/corpora/reference/seg_native_before.json`, which amendment A-01 failed to name.
   Its per-tone shares are reported separately from Lingua Libre's and never pooled with them.
9. **`claimed_tones` is `tones`, the post-sandhi field**, and the 24 covered syllables where
   `tones != citationTones` are published as their own line with both readings.
10. **The five gate outcomes are mapped before the run:** `PASS` is a match or a mismatch according
    to the declared numerator, and `INCONCLUSIVE`, `TONE-FLAG` and `UNUSABLE` are no-verdict and sit
    outside the denominator, with their counts published per tone. The register arm is included.
11. **Uncertainty is reported.** Every share carries a bootstrap interval resampled by speaker and
    by item, not by nucleus, because nuclei within one over-split syllable are not independent.
12. **Every per-speaker figure carries its usable recording count**, and the speaker range is
    computed only over speakers above a recording threshold declared now: **at least 10 usable
    recordings**.

**The re-registration, and why it goes this way.** The adversary showed that expectation 3 as
registered was broken before any audio was read: `voice_median_hz`'s own control puts the tone's
effect on a per-syllable median at 14.09 semitones with the register held fixed, and 84 of the 104
shared items fall below that. The between-speaker term as I specified it was therefore mostly the
tone again.

I am **not** re-registering a corrected range, because the corrected range would be set from the
adversary's numbers, which are numbers about this same sample. Setting a prediction from the data
it predicts is the thing pre-registration exists to prevent, and doing it one step early does not
make it a different act. So:

- **Expectation 3 is withdrawn and becomes `not-predicted`.** The speaker axis is still measured and
  still reported; it simply carries no registered direction, and the article may not claim it was
  predicted. The estimator changes as the adversary asks: a speaker's register is their median over
  all their own syllables on a tone-balanced subset, with the count published, and the between-tone
  term is measured paired inside one recording on the 70 disyllabic items whose two syllables carry
  two distinct non-neutral tones.
- **Expectation 1 stands at 0.10 to 0.40**, with its numerator now declared as `shape_of ==
  DIP-RISE`.
- **Expectation 2 stands**, with the tone-4 numerator declared the same way.

**The replication is redesigned.** The adversary's objection 3.7 is correct: a second session over
the same cached bytes running deterministic code agrees by construction and is evidence of nothing.
The re-execution session now also perturbs every nucleus boundary by the disagreement distance
`segment.seg_confidence_seconds` already computes, and reports how many verdicts move.

### Amendment A-03 — the render disagreed with the registered text, and the analysis restriction

`written_at: 2026-09-10T17:10:00Z` · `post_hoc: true` · `author: the run` · **declared, not repaired**

**A-03.1 The Gate B block that Jakub approved stated a different expectation 2 than §3 of this file.**

Registered in §3: *"Tone 4 is the most reliably shaped tone, matching the notation's fall in at least
0.70 of its nuclei, and tone 1 the second most ... while tone 2 is the least reliable of the four."*

Rendered in `GATE_B.md`: *"Prvý tón bude tvar notácie spĺňať najčastejšie z troch, ktoré majú tvar,
a to aspoň v 0,70 prípadov."* — tone **1** most often, ≥0.70. Expectation 1 was rendered faithfully;
expectation 3 was correctly described as withdrawn.

**The registration governs.** §3 is the artefact; the block was a paraphrase and the paraphrase was
wrong. The measurement is therefore read against §3, and read against the block as a second reading,
and both are published. Measured, on the majority rule:

| tone | notation shape | share matching | interval by item |
|---|---|---|---|
| 1 | LEVEL | 0.8814 | 0.8165 to 0.9364 |
| 2 | RISE | 0.5755 | 0.4711 to 0.6900 |
| 3 | DIP-RISE | 0.2069 | 0.1045 to 0.3273 |
| 4 | FALL | 0.7596 | 0.6700 to 0.8417 |

Read against §3, expectation 2 is **mostly refuted**: tone 4 is not the most reliably shaped tone
(tone 1 is, by 12 points), tone 1 is not second, and tone 2 is not the least reliable of the four
(tone 3 is, by 37 points). Only the "at least 0.70" clause holds, at 0.7596. Read against the block
that was approved, it is **confirmed**. Both readings ship.

This is a defect in the run and not in the measurement, and it is the reason a render must be
generated from the registered file rather than written beside it. Recorded in `PROPOSALS.md`.

**A-03.2 The per-tone figures are computed only on files where the nucleus count equals the claimed
syllable count.** Registered §2 says the denominator is "nuclei that reached a verdict" and does not
say this. The reason is structural rather than statistical: a nucleus is matched to a tone by its
index, and the index means a tone only when the counts agree. On a file with three nuclei and two
claimed syllables there is no way to say which nucleus is which tone without a decision the method
does not license.

Consequence, published rather than buried: the restriction removes 315 of the 624 files, and it does
not remove them evenly. Tone 3 keeps the thinnest denominator of the four, 58 graded nuclei, because
its files are the ones the segmenter handles worst.

Nothing about the choice was informed by the shape outcome, which had not been computed when the
restriction was written; the partial agreement rates had been seen. Marked `post_hoc: true` for that
reason and not because the choice depended on the result.

### 9. Output — the field list of §2, restated so a check can read it

`written_at: 2026-09-10T17:40:00Z` · `post_hoc: false` · `author: the run`

This section adds no promise. It restates §2's "Output file and its shape" paragraph under a
heading, because `scripts/verify-study.mjs` finds that paragraph with
`/^###?\s+[\d.]*\s*Output\b/` and §2 carries it as a bolded line rather than a heading, so the check
that compares promised fields against the fields that were written **skipped this run in silence**.
A check that does not run is indistinguishable from a check that passes. The field names below are
copied out of §2 and not reduced.

Promised on `data/units.jsonl`, one row per nucleus: `run_id`, `unit_id`, `item_id`, `hanzi`,
`pinyin`, `claimed_tones`, `syllable_position`, `corpus`, `speaker`, `license`, `file_sha256`,
`nuclei_in_file`, `claimed_syllables`, `segmentation_agrees`, `nucleus_a`, `nucleus_b`, and per
tracker `{shape, net_st, range_st, t_min, t_max, dur_ms, median_hz, n_voiced}`, plus `verdict_tone`
and `verdict_kind`.

Promised on `data/files.jsonl`, one row per file: `fetch_status`, `riffs_check`, `bytes`.

Where the written data differs from that list, `data/deviations.json` names it: D-03 for the
per-tracker fields, D-04 for the per-nucleus ones. Nothing here is corrected in the frozen text.

### A-04 — why that heading exists, and what it immediately caught

`written_at: 2026-09-10T17:40:00Z` · `post_hoc: false` · `author: the run`
`header_added_at: 2026-09-11T01:30:00Z` — see A-08, which records why the header was not there

The restatement above is a correction of the pre-registration's **shape**, not of its content, and
it is recorded rather than done quietly because a run that edits its own frozen text to make a check
runnable is doing the thing the self-modification rule exists to catch. The test that rule sets is
whether the edit could make a future gate easier to pass. This one cannot: it adds no promise, it
adds a field name to a sentence that already held it, and the check it enables was written before
this run.

What it caught on first run is in D-03 and D-04, and it is not cosmetic: the frozen text promises a
per-tracker feature set and the dataset carries one tracker's full set and two trackers' shapes. That
is a promise the run did not keep, and until the heading existed nothing compared the two.

Recorded in `PROPOSALS.md`: the verifier's Output-section check silently skips any pre-registration
that states its output promises anywhere other than under a heading, and silence reads exactly like
a pass.

### A-05 — §9 put two words in section 2's mouth

`written_at: 2026-09-10T17:55:00Z` · `post_hoc: false` · `author: the run`

§9 says its field names "are copied out of §2 and not reduced". For the `units.jsonl` list that is
true. For the one sentence about `files.jsonl` it is not: §2 wrote *"the fetch status, the RIFF check
and the byte count"*, naming all three in prose and none of them as identifiers, and §9 turned two of
them into the field names `fetch_status` and `riffs_check`. Those names appear nowhere in the frozen
text.

The correction is recorded here rather than made in §9, because an amendment is appended and never
edited. The check reads §9, so §9's list is what it enforces, and the run keeps it: the promise §2
made is real even where §9's spelling of it is not, and `data/files.jsonl` does not carry the RIFF
check or a byte count. `data/deviations.json` D-07 records the gap, under both spellings.

Falsified before trusted: the check went red on exactly these two names the first time it ran, which
is the only evidence that it is looking at anything.

### A-06 — the re-execution, what it moved, and the defect it found in this run's own tool

`written_at: 2026-09-10T22:25:00Z` · `post_hoc: true` · `author: the run, on REPLICATION_A.md`

**What ran.** The re-execution of the whole pipeline by a separate DeepSeek session that never saw
this run's dataset, answer or expectation. Its prompt is `_prompts/R1.md` and its report is
`replication/REPLICATION_A.md`; its own datasets are `replication/replication-R0.jsonl` and
`replication-R1.jsonl`, 1253 rows each, and its code is `replication/run_replication.py`. It
reproduced the raw segmentation exactly: 1212 nuclei over 624 files, the same number this run found.

**1. The boundary perturbation moves a third of the verdicts.** That is the answer to adversary
objection 3.7, and it is the most important number in this file. Moving each nucleus boundary by the
disagreement distance `segment.seg_confidence_seconds` computes, with nothing else changed:

| comparison | moved | of |
|---|---|---|
| the nucleus's verdict | 408 | 1253 |
| the majority shape across three trackers | 405 | 1253 |
| any single tracker's shape | 645 | 1253 |

The perturbation is a median 70 ms on Lingua Libre and 160 ms on AISHELL-3. Every per-tone share in
this run therefore carries an uncertainty the bootstrap intervals do not contain, because the
bootstrap resamples which nuclei are in the sample and this resamples where each nucleus is.

**2. The two implementations disagree, and here is how much.** Both sides computed under this run's
own rules — agreeing files only, tone 5 outside the denominator, the five gate outcomes mapped so
that UNUSABLE, INCONCLUSIVE and TONE-FLAG are no-verdict, a tone matched on a 2-of-3 majority —
restricted to the 282 files both implementations call agreeing:

| tone | this run | replication R0 | replication R1 |
|---|---|---|---|
| 1 | 97/111, 0.8739 | 100/110, 0.9091 | 80/107, 0.7477 |
| 2 | 55/99, 0.5556 | 47/97, 0.4845 | 41/88, 0.4659 |
| 3 | 11/52, 0.2115 | 5/49, 0.1020 | 9/51, 0.1765 |
| 4 | 72/96, 0.7500 | 73/97, 0.7526 | 75/100, 0.7500 |

Tones 1 and 4 replicate to within four points. Tones 2 and 3 do not: seven points and eleven points
apart, which on tone 3 is six nuclei out of about fifty and sits inside the interval this run
publishes for that tone, 0.00 to 0.30 by speaker.

**3. Three further measurements to find out why, and two of them clear this run.** The replication
prompt did not carry deviation G-01 or amendment A-03, so two candidate causes were named. Both were
measured rather than argued (`PREREGISTRATION.md` §2 step 3, `data/sens-gatefloor/`,
`data/reg-trackers/`):

| tone | this run | gate floor for everyone | registered `trackers_of` call |
|---|---|---|---|
| 1 | 0.8814 | 0.9076 | 0.8889 |
| 2 | 0.5755 | 0.5500 | 0.5810 |
| 3 | 0.2069 | 0.2000 | 0.2034 |
| 4 | 0.7596 | 0.7611 | 0.7547 |

The floor moves tone 3 by 0.007. The registered tracker call moves it by 0.0035. Neither explains the
gap, and the residual is six nuclei.

**4. The defect those runs found in this run's own tool.** Section 2 step 3 registers
`metrics.trackers_of(path, x, sr, n, [...], segs)`. `tools/measure.py` does not call it. It
reimplements the three trackers so it can set the F0 floor per speaker, and in doing so calls
`metrics.f0_pyin(path, n)` with no span set. `f0_pyin` despikes against the spans it is given, so
with no spans pyin is filtered once against a file-global pitch reference inside the function and
then filtered a second time by the caller against the syllable-local one, and a frame the
file-global pass zeroes cannot be recovered by the second. Pyworld and Praat are filtered once.
The registered call is a new deviation, **G-09**, and its measured effect on the headline is 0.35
percentage points, which is why this run's numbers stand rather than being replaced.

**5. G-01 does not do what its own record says.** G-01 raises the floor to 75 per cent of the
speaker's own tenth percentile and justifies itself with Jouketou, "whose tone 3 bottom falls where
the gate cannot separate it from an error". For Jouketou the formula returns exactly 60.0 Hz, the
gate's own value: `max(60, 76.8 x 0.75) = 60`. Jouketou is the one Lingua Libre speaker of twelve
whose floor the deviation leaves unchanged, and the eleven it does raise are all higher voices. The
deviation is kept because it is declared and because the sensitivity run shows it moves the answer by
less than a point, but its stated reason is wrong about its own example and a reader must not be
given it.

**6. Disposition of the ledger.** Tone 3 rows are `disagreed`. Tones 1, 2 and 4 are `re-executed`,
with the replication's reading published beside this run's in every case, and the residual difference
named. The article states plainly that the third tone's share did not replicate to within the design's
own precision while its direction did, and it may not quote a single tone-3 figure as though the
replication did not exist.

**7. One more defect, this one in the replication prompt.** The prompt handed the replicator §2 as
amended by A-02 and left out G-01 and A-03. A re-execution is handed the method that was actually
run, deviations included; anything left out of that prompt becomes a disagreement that has to be
chased afterwards, at the cost of two extra measurements. This is INC-shaped and belongs in
`PROPOSALS.md`.

### A-07 — the independent derivation, and the one thing it settles and this run cannot

`written_at: 2026-09-10T23:05:00Z` · `post_hoc: true` · `author: the run, on REPLICATION_B.md`

**What ran.** A second session was given the question and the two corpora and nothing else: not this
run's method, not its dataset, not its answer, not the adversary's report. It wrote its own method
down first (`replication/METHOD.md`), recorded seven later changes with their reasons
(`replication/METHOD-AMENDMENTS.md`), and answered. Its numbers are in
`data/derivation-results.json` and in `replication/REPLICATION_B.md`.

**1. The two methods put the tones in different orders, and that is a result.**

| tone | this run | the derivation |
|---|---|---|
| first, level | 0.8814 (118) | 0.579 (292) |
| second, rising | 0.5755 (106) | 0.521 (282) |
| third, dip | 0.2069 (58) | 0.135 (192) |
| fourth, falling | 0.7596 (104) | 0.781 (319) |

The third tone is last by a wide margin in both, and the second is in the lower half in both. The
first and fourth tones trade places. The reason is not a defect in either: the two methods define
"level" differently. This run asks whether the shape a tracker reports is `LEVEL` on a
syllable-normalised contour; the derivation smooths the contour, reads it at ten points and calls it
level when it moves less than half a Chao level on the speaker's own grid. A strict flatness test
costs the first tone almost thirty points and costs the fourth tone almost nothing, because a fall of
three levels is unmistakable and a level contour is a claim about how little something moved.

**Which means expectation 2 is not refuted, it is undecided.** This run's own reading refutes the
registered text (tone 1 matches most often, tone 4 second). The derivation reproduces the registered
ordering (tone 4 first, tone 1 second). *Which tone's notation shape is realised most reliably depends
on how strictly "level" is drawn*, and the article must say that rather than picking the reading it
likes. The registration named tone 4 and the registration is not vindicated by one of two competent
methods agreeing with it after the fact.

**2. The derivation carried the check this run could not, and reached the opposite of a
reassurance.** Its own largest stated limitation is that the third tone's rise happens in creaky
voice, where pitch trackers stop, so the window it measures is missing the very part of the syllable
the notation describes. It tested this by splitting third-tone syllables into quartiles of window
coverage: the dipping rate climbs 2.1 per cent, 12.5, 20.8, 18.8. Some of the low third-tone figure is
therefore its own instrument, and it says so.

**3. This run ran the same check on its own data and could not carry it** (`C-96`, `C-97`,
`data/coverage-gradient.json`, `tools/analyse-coverage.py`, written after reading the derivation's
result and therefore post_hoc). Its third-tone quartiles hold fourteen nuclei each, which puts an
interval of about a quarter of a share on every figure: the shares come out 0.2667, 0.1333, 0.0714,
0.3571 with no usable gradient. The fourth tone, the control both designs used, does show the
gradient this run can read, 0.6538 to 0.8462, and a rising gradient on a falling tone is what a
truncated window produces rather than anything about the notation.

**4. What this changes in the article.** Three things.

- The tone-3 share is reported with the tracker-coverage limitation attached at the point of the
  claim, not in a closing caveat, and the article states that neither design can separate a real
  low-falling allotone from a truncated rise. The derivation's own numbers for that shape, 0.74 of
  its third-tone syllables measuring as a plain fall, are quoted beside it.
- The comparison between the first and fourth tones is written as unsettled and the two methods'
  figures are published side by side.
- The article's title may not rest on which tone matches most often. It rests on the third tone,
  where both methods agree, and on the fact that the answer is a number a reader can check.

### A-08 — A-04 was written without its header, and the header was added afterwards

`written_at: 2026-09-11T01:30:00Z` · `post_hoc: false` · `author: the run`

A-04 was appended at 17:40 on 2026-09-10 with its heading and its argument and without the
`written_at` and `post_hoc` line that every other amendment carries. Nothing noticed for five hours,
including the session that wrote it and the two sessions that revised the article afterwards.

The run's own mechanical compliance check found it, on the first execution, by reading the
pre-registration for amendment headings and requiring the two fields under each. That is the check
catching a real defect on its first run, which is also the only evidence that it looks at anything.

**The header was inserted and nothing else was touched.** A-04's argument is unchanged, character for
character. The insertion is recorded here rather than passed over because an amendment that is
edited after the fact is exactly what the append-only rule exists to make visible, and an invisible
insertion would be worse than the missing header it repairs. The header carries the time A-04 was
actually written, not the time the header was added, and the added line carries both.

**Proposal.** The amendment header should be checked at the moment the amendment is appended, by
whatever writes the amendment, rather than hours later by the compliance pass. `PROPOSALS.md` carries
it as item 13.

### A-09 — the re-execution disagreement is resolved, and it was one ambiguous sentence

`written_at: 2026-09-11T02:05:00+07:00` · `post_hoc: true` · `author: the run, on an audit finding`

**What was open.** A-06 recorded that the re-execution disagreed with this run on the second and
third tones and left the rows marked `disagreed`, on the reasoning that both readings would ship. An
audit of the finished run found that wrong on the skill's own terms: `replication.md` treats a
re-execution disagreement as a defect that blocks the affected rows until the disagreement is
resolved, while publishing both readings is what it asks for a *refuted expectation*, which is a
finding about the world rather than a fault in the instrument.

**What the disagreement was.** Not the data. `tools/diagnose-disagreement.py` takes one of the five
tone-3 nuclei the two implementations grade differently — `Jouketou-百.wav`, the lowest voice in the
corpus — and runs both pitch paths over it. The two produce **identical F0 tracks**: 119, 109 and 100
voiced frames for pyworld, Praat and pyin on both sides, differing on zero frames. What differs is the
window.

The word is one syllable. `segment_energy` splits it into two nuclei, `[8, 95]` and `[95, 121]`, and
`segment.artefact_nuclei` marks the second as an artefact of the split. This run merges the artefact
back and grades the syllable as `[8, 121]`, where all three trackers read `DIP-RISE`. The replication
grades the first nucleus as the segmenter left it, `[8, 95]`, where two of three read `FALL`. The
replication's own file row carries `artefact_nuclei: [1]` and `nuclei_corrected: 1`, so it found the
artefact and applied it to the agreement count and not to the grading, and it recorded that choice as
an ambiguity at the time: *"Provision E's 'artefact_nuclei applied before grading' against the
discard rule: I kept every nucleus graded … and applied artefact_nuclei only to the over-split count
in the coverage statistics."*

So the fault is this run's, and it is a fault in **the written description**, not in the measurement.
Provision E as it was handed over — *"`segment.artefact_nuclei` is applied to the over-split nuclei
before grading"* — does not say that grading uses the merged span, and the replication read it the
other way. That is the outcome `replication.md` predicts for this pass: *"A step whose written
description does not actually reproduce what the first session did, which is the common case and the
most useful thing this pass finds."*

**The reconciliation, measured rather than argued.** If the explanation is right, one switch in this
run's own pipeline should reproduce the re-execution's number. `tools/measure-nomerge.py` is the same
code with the artefact merge switched off, and it is the only difference.

| tone | this run | the same code with no merge | the re-execution, on the files both call agreeing |
|---|---|---|---|
| first | 0.8814 (118) | 0.8462 (195) | 0.9091 (110) |
| second | 0.5755 (106) | 0.4379 (169) | 0.4845 (97) |
| third | 0.2069 (58) | **0.0988 (81)** | **0.1020 (49)** |
| fourth | 0.7596 (104) | 0.7665 (167) | 0.7526 (97) |

Tone 3 lands three thousandths from the re-execution's figure — 0.0988 against 0.1020 — and the
re-execution's tone 3 was the whole of the disagreement that mattered. The remaining spread on tones
1 and 2 is the difference in which files the two call agreeing, 399 against 295, which follows from
the same switch.

**Disposition.** The tone-3 rows stop being `disagreed` and become `re-executed`, with the ambiguity
named. Nothing about the numbers changes: this run keeps the merge, because the merge is what
recovers the syllable the segmenter split, and the file inspected is a word of one syllable whose
true span is the merged one. **What the article must now carry is the switch itself**, because
whether an over-split nucleus is graded where the segmenter put it or merged back into its syllable
moves the headline from 0.21 to 0.10 — a larger difference than any interval this run publishes, and
the second-largest sensitivity after the boundary move.

**A correction to A-06.** A-06 said the residual was "six nuclei out of about fifty" and called it a
small difference inside the published interval. Both true, and both beside the point: the residual had
a single identifiable cause, and calling it noise was the wrong call. It was found by an audit of the
finished run, four hours after A-06 was written, and not by this run.

### A-10 — the timestamps, and the ordering they were supposed to establish

`written_at: 2026-09-11T02:20:00Z` · `post_hoc: true` · `author: the run, on an audit finding`

**What is wrong.** Nine `written_at` stamps sit above. They are the machine's local clock, which is
UTC+7, with a `Z` attached. Read as UTC they are wrong by seven hours, and A-06, A-07 and A-08 read
as being in the future when they were written; read as local they are right. The file therefore
mixes two time bases and no reader can tell which.

**Why they are not edited.** A-08 exists because an amendment was edited after the fact, and it says
in its own words that an invisible insertion would be worse than the missing header it repaired.
Replacing nine stamps in place would be that failure nine times over. They stay as they are and the
true ordering is established here, from file birth times and from the child sessions' own logs,
which no one typed.

**The ordering, in UTC, from evidence.**

| Time (UTC) | What happened | Evidence |
|---|---|---|
| 12:47:12 | `GATE_A.md` written, seven candidates, no demand line | file birth time |
| 13:02:47 | Jakub: *"Ok spravme B"* | the session transcript |
| 13:07:29 to 13:10:57 | the demand pass runs | `_logs/C1.jsonl`, `_logs/C2.jsonl` |
| 13:26 | Jakub asks for the numbers before the choice | the session transcript |
| 13:29:39 | `STEP3_CANDIDATE_NUMBERS.md` written, the same minute it was posted | file birth time |
| 13:44:54 | Jakub: *"ok podme teda pokracovat s temou A"* | the session transcript |
| 14:34:48 | `PREREGISTRATION.md` created, **before the adversary is launched** | file birth time |
| 14:35:00 to 14:56:13 | the adversary runs | `_logs/A1.jsonl` |
| **14:53:41** | **`GATE_B.md` written, carrying the 21 dispositions** | file birth time |
| **14:56:18** | **the first measurement row** | `data/files.jsonl`, `measured_at` |
| 15:00:31 | the last measurement row | `data/files.jsonl` |
| 15:03:06 to 15:09:31 | re-execution | `_logs/R1.jsonl` |
| 15:03:08 to 15:42:41 | independent derivation | `_logs/R2.jsonl` |
| 15:16:49 to 16:54:32 | the writers, in three passes | `_logs/W1..W3.jsonl` |

**The ordering the file exists to prove.** The adversary's 21 objections were dispositioned into the
method **before the first file was opened**. The evidence is `GATE_B.md`'s birth time, 14:53:41Z,
which is three minutes before the first measurement row at 14:56:18Z, and `GATE_B.md` exists for no
other purpose than to carry those dispositions to Jakub. A-02's own stamp cannot be used for this and
is not used: it is written in the local time base, and the file that would settle it says so.

**What this does not fix.** A-02's stamp still reads as a time at which the measurement was already
running. Nothing in the run can repair that without editing a frozen-section amendment, and editing
it is the thing this file's own A-08 forbids. The stamp is wrong and the ordering is right, and both
statements are now on the record.

### A-11 — the dataset was re-measured to carry the schema section 2 registered, and the pinyin was wrong

`written_at: 2026-09-11T02:30:00Z` · `post_hoc: true` · `author: the run, on an audit finding`

**What the audit found.** Section 2 registers the fields `data/units.jsonl` carries, and Rule 2 of
`protocol.md` requires six provenance fields on every row. The file that shipped had neither. Its
complete key set was 22 names, none of them `measured_at`, `input_sha256`, `tool`, `tool_version`,
`command`, `log`, `expected` or `outcome`; ten of the registered fields were absent, including
`item_id`, `hanzi`, `syllable_position`, `license`, `file_sha256`, `nuclei_in_file`,
`claimed_syllables`, `segmentation_agrees`, `verdict_tone` and `verdict_kind`; and the registered
per-tracker block survived for pyworld alone, so the three-tracker agreement the method sells could
not be re-derived from the published data at all. The severity was `BLOCKS_PUBLICATION`, because a
Standard-tier run's dataset is the thing a reader is invited to check.

**What was done.** `tools/measure-registered.py` is the same pipeline with the row schema completed,
and it was re-run over the same 624 files under the same environment. Both passes are on disk:
`data/units-pass1.jsonl` and `data/files-pass1.jsonl` are the first, `data/units.jsonl` and
`data/files.jsonl` are the registered one, and the tables, the ledger and the analysis chain were
rebuilt from the second.

**The re-run reproduces the first pass exactly.** On all 975 nuclei and on every field the two share
— `nucleus_a`, `nucleus_b`, `window_a`, `window_b`, `shape_by_tracker`, `satisfied_union`,
`gate_verdict`, `raw_nuclei_merged` — the number of differences is **zero**. Every analysis output
is byte-identical except the two things the repair changed on purpose. A run that re-measures its
own corpus and gets the same answer is the weakest useful replication, and it is the one that was
available.

**A defect this repair found on its own.** The Lingua Libre pinyin in `corpus.generated.ts` is a
space-separated string, not a list, and the first pass iterated it as a string. `onset_class` was
therefore computed once per **character**: `yì xiē` produced six onset classes instead of two, and
`onset_classes` in the published dataset was meaningless. No article number is split by onset class
— the ledger contains no row that selects one — so nothing printed moved, and the defect was confined
to a published column nobody had used. It is repaired in the registered pass and it is recorded here
rather than quietly overwritten, because the published dataset is the thing a reader checks and it
was wrong for an afternoon.

**What the numbers did.** Nothing in the article changed. The two intentional differences between the
passes are the onset-class counts and the pinyin field's type; every other analysis output is
identical, and the ledger was rebuilt and re-verified against the regenerated datasets with no row
moving.

**A-09's switch belongs to the emitted method, not to the data.** No sentence in this file or in
`METHOD.md` said which span a merged syllable is graded on, and the re-execution's reading of the one
that tried is the one the replicator recorded. The description is now unambiguous in the re-execution
prompt (`_prompts/R1b.md`) and the reading is named in A-09.

### A-12 — the re-execution re-run with the clarified sentence, and a correction to A-09

`written_at: 2026-09-11T02:40:00Z` · `post_hoc: true` · `author: the run, on the re-execution it re-ran`

**What was re-run.** `replication.md` asks for both sides to be re-run before a disagreed row reaches
the article. This run's side did not change — the merge is what it always did — and the re-execution
was launched again with provision E rewritten to say which span is graded, everything else identical
(`_prompts/R1b.md`, `_logs/R1b.jsonl`, `rerun/REPLICATION_A2.md`). Same model, same corpora, same
reporting.

**What it found.** Its Lingua Libre shares are now first tone 0.8095, second 0.5676, **third 0.1688**
and fourth 0.7578. The first re-execution had the third tone at 0.1020 and the gap was eleven points;
with the sentence made explicit it is four. The clarification was therefore the cause of most of the
disagreement, which is what A-09 predicted.

**Where A-09 overreached, and this corrects it.** A-09 said the disagreement was "fully explained by
the segment merge" and that one switch "reconciles the two". That is too strong on two counts, and
the corrected re-execution is what shows it.

- **It does not reproduce this run's figure.** 0.1688 on Lingua Libre against this run's 0.2069 over
  the agreeing files of both corpora. The denominators differ — all nuclei that reached a verdict on
  Lingua Libre, against nuclei on files where the segmentation agrees, pooled — so the two are
  comparable in order and not in level. That is the same caveat the article attaches to the
  independent derivation, and it applies here too.
- **It cannot reproduce the 0.10 at all.** It tried four span choices and none gives 0.21, and 0.10
  appears twice without its partner. So the pair A-09 quoted as reconciled is a pair that only this
  run's own pipeline produces. A-09's table is sound — the same code with the merge off really does
  give 0.0988 — but the claim that this reconciles *both sides* was this run reading its own result
  as the other side's.

**A finding about this run's instrument, which the re-execution is right about and which is verified
here.** `segment.artefact_nuclei`'s own docstring calls it a test of whether an over-count *can be
explained*. This run applies it wherever it names an index, and the merge then fires on files that
were never over-split. Counted directly on `data/files.jsonl`: the number of files whose nucleus
count equals the claimed syllable count falls from **399 of 624** on the raw count to **309** after
the merge, and on AISHELL-3 alone from **108 to 34**. Seventy-four AISHELL files that the segmenter
had already cut correctly are merged anyway. That is the widest spread in the article's table and
this is why, and it belongs in the article rather than in this file: the third-tone figure's largest
single uncertainty is a rule of this run's own, applied more widely than the function's docstring
describes.

**What it changes.** The article's tone-3 rows keep both readings, and one sentence is added to the
paragraph that carries the switch, saying the merge fires where there is no over-count and what that
costs the AISHELL-3 files. No published number changes. The verdict on the disagreement is
"resolved as to its cause and reproducible in the order, not in the level".
