File the J4 word round, and read its broken tie gate class by class - #554
Merged
Conversation
Round 5 of the human authenticity pass is judged (2026-09-06, 75 screens, one sitting). The candidate — the J4 exit trim, `exit_trim` — takes 34 of the 36 decided screens (94.4 % against a >= 60 % threshold), while the pre-registered tie threshold of <= 25 % fails at 42.9 %. `adopt: false` stands as the tool reported it. The contradiction dissolves per class, and the classes were declared before the round: `naht-stark` meets BOTH thresholds (26 : 2, 9.7 % ties), `naht-schwach` runs 8 : 0 for the candidate at 72.4 % ties. That is the class the pre-registration had described as one where the trim probably does not show. 21 of the 27 ties come from it; the control class is a clean 3 of 3. Re-measured on today's frozen sep05 root, the arm costs +0.000581 on words, leaves the pairs byte-identical and drops the seam departure from +7.59 to -0.70 degrees (absolute median 12.67 -> 2.30), doublings 14 = 14. The whole ruler loss sits in `naht-stark` (+0.036888 against -0.000262), which is exactly where the eye votes 26 : 2 FOR the trim; of the 30 words the ruler punishes, 18 go to the candidate and none to the base. The obvious narrowing was measured, not argued: the existing `exit_trim_min_kink_deg` knob separates neither class at any rung — both carry the same kink, +7.90 against +7.44 — costs more than the full trim between 5 and 25 degrees, and gives the seam repair back. So the global flip is the smaller honest mechanism, and the entry recommends it; the switch itself is a rendering-affecting change and therefore the author's call, as the LF11 round was. Filed like round 6: judgements, narrow key, analysis JSON and a provenance stamp that also carries the dirty-worktree flag of the build commit. The full key, the payload, the arm files and the strata file stay out. `menschliche-bewertung.md` gains the instrument note as an explicit PROPOSAL: a tie threshold defined over all screens also measures the class mix a round chose, and a threshold that follows a round's own class declaration can be softened by cutting the classes differently — which is why it is not adopted here. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
There was a problem hiding this comment.
🟡 Changes recommended
Provenance and reconstruction claims are inaccurate, and parts of the class analysis contradict their own measurements.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Files and interprets humanbench round 5 for J4 without changing rendering behavior.
Changes:
- Archives judgments, slim key, analysis, and provenance.
- Documents class-level results and the tie-gate proposal.
- Updates J4 tracking and changelog records.
File summaries
| File | Description |
|---|---|
docs/reference/messjournal.md |
Records round results and analysis. |
docs/reference/menschliche-bewertung.md |
Proposes class-level tie-gate handling. |
docs/proposals/tintenfolger.md |
Updates J4 rescue-path status. |
data/humanbench/SOURCE.md |
Adds round 5 provenance. |
data/humanbench/runde-05-vorkommen.json |
Stores the slim round key. |
data/humanbench/runde-05-urteile.txt |
Stores judgments. |
data/humanbench/runde-05-stempel.md |
Documents build provenance. |
data/humanbench/runde-05-auswertung.json |
Stores aggregate analysis. |
data/DATA_PROVENANCE.md |
Indexes round 5 artifacts. |
changelog.d/runde-05-j4-ablage.md |
Adds the changelog fragment. |
Review details
Suppressed comments (4)
data/humanbench/runde-05-stempel.md:41
check_arm_scopedoes not verify this digest. It comparesstyle,source_id,fixture_root, and the arms'exported_atvalues (tools/humanbench/build.py:1159-1197);wordarm.pydoes not store a root digest in either arm. Please distinguish the manually recorded digest from the fields the builder actually checked.
| Wurzel-Export | `exported_at 2026-09-02T22:16:06+00:00`, Digest `6cbab9d5c092` (beide Arme, gegeneinander geprüft von `build.py::check_arm_scope`) |
data/humanbench/runde-05-stempel.md:184
- This deterministic-rebuild claim is not supported while the recorded build has
code_dirty: true: the stamp omits the dirty diff, so the seed, root, and commit do not identify the builder code that produced the omitted full key/payload. This should either point to an archived dirty diff or explicitly state that exact reconstruction is unavailable.
Sie bleiben unter `temp/runden-sep04/humanbench/` und sind aus Saat, Wurzel und
diesem Stempel deterministisch wiederherstellbar; die Klassenzuordnung selbst
steht Wort für Wort im schmalen Schlüssel (`stratum`).
data/humanbench/runde-05-stempel.md:199
- These reconstruction commands omit
--fixtures, so they read whichever export currently occupies the default fixture path. The stamp already says the judged round used6cbab9d5c092while today's root iseaa195aa7c84; running these commands now therefore creates different arms and a different round. Point bothwordarmcalls andbuildat a preserved parent containing the September 2 export, then verify the recorded arm hashes.
uv run python -m tools.humanbench.wordarm --arm "Basis (LF11, Chart-Nib)" \
--out temp/runden-sep04/humanbench/arm-basis.json
uv run python -m tools.humanbench.wordarm --arm "J4 Austritts-Trim" --exit-trim \
--registration-from temp/runden-sep04/humanbench/arm-basis.json \
--out temp/runden-sep04/humanbench/arm-j4.json
data/humanbench/runde-05-stempel.md:120
- The claim that the dirty tree is harmless contradicts the method's own provenance rule:
menschliche-bewertung.md:879-885says that withcode_dirty: true, the commit is only a clue, not proof of the code used. Arm hashes pin geometry, but not dirty builder or strata changes that determine selection, order, and mirroring. Record what was dirty before claiming the commit determines those choices.
> **Der unsaubere Arbeitsbaum gehört genannt, und er ist hier folgenlos.** Was
> die Runde zeigt, sind zwei fertige Arm-DATEIEN mit ihren `sha256`; die Seite
> komponiert nichts nach. Der Bau-Commit bestimmt also die Auswahl, die
> Anordnung und die Spiegelung — nicht die Geometrie. Trotzdem steht das
> Flag hier: ein Stempel, der nur die bequemen Felder trägt, ist keiner.
- Files reviewed: 10/10 changed files
- Comments generated: 5
- Review effort level: Balanced
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…ot separate Five findings from the Copilot pass, four of them substantive. The ladder analysis overstated its own table. `exit_trim_min_kink_deg` DOES enrich `naht-stark` from 10 degrees on — 71 % of it still firing against 59 % of the weak class, 39 % against 14 % at 25 — and at 5 degrees the enrichment even runs the other way. What it never does is SEPARATE the two, which is the narrower and defensible claim: at 20 degrees it still fires in seven weak words while already dropping nineteen strong ones, and the reason is the measurement above it, that both classes carry the same seam kink. The second overstatement was the closing sentence: 30 degrees does pay 0.000202 of the 0.000581 back, so the ruler rewards exactly one rung — it just buys that third of the price by giving up 55 of the 60 words the trim fires on, four strong and one weak surviving. The provenance note read as if only rounds 4 and 5 were built on September 4. All three were; 6 was judged on the 5th, 5 on the 6th, and 4 is still unjudged. Said that way in the stamp and in SOURCE.md. The §7.11 row claimed three of the four J4 conversions were done. Two are: `dspan` and the word round. The arrival side and the Endblenden path are both open, and the row now says so, which is what §7.9 said all along. And §8a's Stand block still read "zweimal gefahren" with only the LF11 and J5 rounds in it. It now carries the third and what it did to the tie threshold. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
Owner
Author
|
Review answered in 63a149a — four substantive findings, all of them right.
|
…uffix The sentence promised a "Nachtrag" while the entry that books it is a section of its own — a small thing, but the register is read by its shape, and a pointer that names the wrong shape sends the next reader looking in this section instead of the one below it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
"Am Lineal teurer" was true of five of the narrowing's six rungs and false of the sixth — the same overreach the review caught two paragraphs above, one sentence later. Say five of six, and say what the sixth does instead. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
# Conflicts: # docs/reference/messjournal.md
# Conflicts: # docs/reference/messjournal.md
MarkusNeusinger
added a commit
that referenced
this pull request
Sep 6, 2026
**The owner approved this on 2026-09-06 ("1 ja") — booked as author
decision A37.** Stacked on #554, which files the round this decision
rests on; **merge #554 first**, this PR's base is its branch.
Turns `exit_trim` into the production default. The switch stays — the
pre-adoption base is a bench arm like any other — but production now
writes on the other side of it, and every `/write/word` answer moves
with it (the edge cache holds the old one for up to 24 h, as after the
LF11 write). No DB row is touched: this decision flips a rule, it does
not write geometry.
## The re-baseline, on an UNCHANGED root
`--expect-root eaa195aa7c84,0fbde2d72b64`, `OPENBLAS_NUM_THREADS=1
OMP_NUM_THREADS=1`. The root does not move, so these two lines are
paired, not merely consecutive:
| | words | pairs | `seam_dep_median` | abs. median | `seam_arr_median`
| `gleichzug_doublings` |
|---|---|---|---|---|---|---|
| before (`--no-exit-trim`) | 0.108444 | 0.148236 | +7.59° | 12.67 |
−2.83° | 14 |
| **after (default)** | **0.109026** | **0.148236** | **−0.70°** |
**2.30** | −3.87° | 14 |
| Δ | **+0.000581** | **0.000000** | −8.29° | −10.37 | −1.04° | 0 |
63/63 words and 33/33 pair drills scored, none skipped, none failed.
`worst_word` moves `han` 0.232609 → `regieren` 0.233052; `worst_pair`
stays `In` 0.283819. Components: `comp_transition` 0.089804 → 0.091028,
`comp_coverage` 0.101580 → 0.101667, `comp_width` 0.162397 unchanged —
the trim touches the seam, not the width. The doublings at the delivered
nib stay at 14, which was not a given for a rule that takes ink away.
**The word ruler rises knowingly.** Round 5 showed the whole of that
cost sits in `naht-stark`, the class where the eye votes 26 : 2 **for**
the trim; of the 30 words the ruler punishes, 18 go to the candidate at
the eye and none to the base. `EXIT_TRIM_MIN_KINK_DEG` stays 0.0 — the
same round measured the narrowing and it separates neither class at any
rung.
## The golden fixture: a declared re-baseline
Re-baked with `REGEN_GOLDEN=1`. Measured before the regen:
- **10 of the 11 words move**, `wovon` does not (its exits are backward,
bow and arm exits, which the class excludes by construction).
- **No word gains or loses a draw item** — item counts 6 … 18 before and
after; 2 to 8 items move per word, namely the trimmed letters and their
connectors.
- **A solitary glyph with `pen=None` stays byte-identical** — checked
across all 23 glyphs in the golden payloads, 0 move, because the rule
cannot fire without a following connector. The CLAUDE.md invariant holds
as written.
## Two measurement layers had to follow, and one is a repair
**`pairlab.prodconn` had named this exact case itself.** Its docstring
said the trim replaces the generator's return value inside
`compose_word`, "harmless while the switch is off — should it ever
become the default, this function has to grow the same post-processing
or the dissection will quietly measure the wrong curve." It now does.
The recorder additionally captures `_cut_exit_stub`, which runs if and
only if the trim was really applied (guards, collinear cut and the
min-kink narrowing all sit in front of it), so what is recorded is the
DECISION rather than a re-derivation of it; `replay` shifts that stub
with its letter and calls `_exit_trim_index`/`_straight_to` again. The
rule is never restated in `tools/`. The parity tests now say it in two
halves: an untrimmed join reproduces the recorded call point for point,
a trimmed one reproduces the rule's own two invariants — the coupling
point does not move, and the join is a straight line.
**The chain is exempt, and that is measured rather than assumed.** Its
init default is the frozen mirror (the production init was measured as
K-F on `sep04` and rejected), and its guard reads the composition soll.
`ductus_soll` over all 63 word samples returns **126 rows of which 0
move**: crossings 292, zones 177, strokes 111, touches 89, overlaps 11
on both sides. The trim cuts a stub and straightens a connector; it does
not touch the composed word's topology, so `k0eval`, the structure guard
and every chain number stand on unchanged ground.
**What is left due, named rather than done here:** the pilot's map IS
the composed path, so its `sep05` dev-19 numbers now stand on a
composition that no longer exists. That is a route measurement with its
own protocol (`/verify-trace`), not something an adoption PR should fold
in — it is filed as an open arm in `tintenfolger.md` §7.11 and is due
before the next pilot statement.
## Docs and gates
`messjournal.md` gains the dated adoption entry plus its register row
and a headline-ledger row; `qualitaetsmetrik.md`'s status block carries
the new headline (and lost two lines to stay under its own 40-line cap);
the glossary entry for Austritts-Trim flips from "opt-in, not adopted"
to adopted with both instruments' numbers; `werkzeuge.md` and
`write-api.md` follow.
`mess-runde` is raised to 23 567 — measured 21 425 plus the documented
10 %. The reason is the one the budget's own comment already licenses:
the register books the standard PAIR of rows for a round and the
adoption it triggered. Both rows were condensed to roughly 300 tokens
each (they stood near 500, three times the register's ~163-token
average) and the §7.11 J4 row was rewritten SHORTER than it was, so the
residual prose growth is one new open-arm row and 11 tokens in a Stand
block.
Gates green locally: `pytest` 2478 passed / 8 skipped, `pre-commit run
--all-files`, `tools.changelog check`, `tools.docs_register check` (2
§14 entries, register agrees), `tools.docs_budget check`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Round 5 of the human authenticity pass is judged (2026-09-06, author, 75 screens in one sitting) and filed the way round 6 was. The candidate is the J4 exit trim —
exit_trim, one switch, placement pinned. This PR files and reads the round; it does not flip anything. The adoption is a separate PR.The result
reliable— judged by the picture, not the positionadopt: falseThe tie gate breaks in the class that was predicted to break it
The classes were declared before the round, cut on how far the candidate moves the drawing:
naht-stark(Δ >= 0.1186 xh)naht-schwachunberuehrt(rule does not fire)MIN_PAIRED_PER_CLASS21 of the 27 ties come from
naht-schwach— the class for which the pre-registration wrote that the trim is where it shows if it shows anywhere, i.e. the class where a difference was not expected to be visible. It is not neutral against the arm either: of its eight decided screens, eight go to the candidate and none to the base. The control class is a clean 3 of 3, and twelve mirrored repeats name the same arm ten times at only four same sides — the instrument sees.Ruler and eye point in opposite directions, word for word
Re-measured on today's frozen sep05 root (
--expect-root eaa195aa7c84,0fbde2d72b64, BLAS pinned):seam_dep_mediangleichzug_doublings--exit-trim)Decomposed by the same classes:
word_lossseam_depbase → armnaht-starknaht-schwachThe whole ruler loss sits in the class where the eye votes 26 : 2 for the trim. Of the 30 words the ruler punishes, 18 go to the candidate at the eye and none to the base; the two base votes of the entire round fall on words the ruler credits to the candidate. The seam repair is the same size in both classes (+7.90 vs +7.44) — they differ in displacement, not in kink.
Global flip or class rule? Measured, not argued
The narrowing is not an idea but an existing knob (
exit_trim_min_kink_deg, the J4b arm). The ladder on the same root:naht-starknaht-schwachword_lossseam_dep_medianNo rung separates the classes — it cannot, because their kink is the same. Every rung between 5° and 25° costs MORE than the full trim, and from 10° on the seam repair is handed back. The global flip is the smaller honest mechanism, and trimming the weak class along with it costs nothing: there the ruler is a wash and the eye is 8 : 0 for the trim.
What this PR does and does not claim
adopt: falsestands as the tool reported it — the plan needs both thresholds. What the pre-registration licenses for a result >= 60 % is putting it to the author, which is what the entry does. The precedent is the LF11 word round: direction overwhelming, tie threshold failing in every reading, no formal claim, and the author released on that basis. The entry's recommendation isexit_trimas the production default, as a global flip; the switch itself is a rendering-affecting change and therefore the author's call.Filed like round 6: judgements, narrow key (uid → entry, text, class,
repeat_of), analysis JSON and a provenance stamp that also carries the build commit's dirty-worktree flag. The full key, the payload, both arm files and the strata file with its per-word displacements stay out.menschliche-bewertung.mdgains the instrument note as an explicit proposal, not a rule change: a tie threshold defined over all screens also measures the class mix a round chose — and a threshold that follows a round's own class declaration can be softened by cutting the classes differently, which is why it is not adopted here.Gates
tools.changelog check,tools.docs_register check(1 §14 entry, register agrees),tools.docs_budget check(mess-runde20 943 of 21 057 — the register grew by its one row and §7.11 was condensed to pay for it, so no budget was raised),pre-commit run --all-files,pytest2475 passed. Licence side: four text files under an existing, indexed source with a completeSOURCE.md, no binaries, no payloads,tests/test_reserved_history.pygreen.🤖 Generated with Claude Code
https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3