Skip to content

File the J4 word round, and read its broken tie gate class by class - #554

Merged
MarkusNeusinger merged 6 commits into
mainfrom
runde-05-j4-ablage
Sep 6, 2026
Merged

File the J4 word round, and read its broken tie gate class by class#554
MarkusNeusinger merged 6 commits into
mainfrom
runde-05-j4-ablage

Conversation

@MarkusNeusinger

Copy link
Copy Markdown
Owner

Round 5 of the human authenticity pass is judged (2026-09-06, author, 75 screens in one sitting) and filed the way round 6 was. The candidate is the J4 exit trim — exit_trim, one switch, placement pinned. This PR files and reads the round; it does not flip anything. The adoption is a separate PR.

The result

Step Number Threshold Reading
1 Reliability 10/12 mirrored pairs same arm, 4/12 same side >= 6 pairs, > 7/12 arm both taken, band reliable — judged by the picture, not the position
2 Side balance left 17 · right 19 · ties 27 reported only as even as a side balance gets
3 Verdict candidate 34 : base 2 of 36 decided → 94.4 %; ties 27 of 63 → 42.9 % >= 60 % · <= 25 % candidate threshold beaten by 34 points, tie threshold missed by 18 → adopt: false

The tie gate breaks in the class that was predicted to break it

The classes were declared before the round, cut on how far the candidate moves the drawing:

Class n decided candidate ties thresholds
naht-stark (Δ >= 0.1186 xh) 31 28 26 (92.9 %) 3 (9.7 %) both met
naht-schwach 29 8 8 (100 %) 21 (72.4 %) candidate yes, ties no
unberuehrt (rule does not fire) 3 0 3 (100 %) control, below MIN_PAIRED_PER_CLASS

21 of the 27 ties come from naht-schwach — the class for which the pre-registration wrote that the trim is where it shows if it shows anywhere, i.e. the class where a difference was not expected to be visible. It is not neutral against the arm either: of its eight decided screens, eight go to the candidate and none to the base. The control class is a clean 3 of 3, and twelve mirrored repeats name the same arm ten times at only four same sides — the instrument sees.

Ruler and eye point in opposite directions, word for word

Re-measured on today's frozen sep05 root (--expect-root eaa195aa7c84,0fbde2d72b64, BLAS pinned):

words pairs seam_dep_median abs. median gleichzug_doublings
base 0.108444 0.148236 +7.59° 12.67 14
J4 (--exit-trim) 0.109026 (+0.000581) 0.148236 (byte-identical) −0.70° 2.30 14

Decomposed by the same classes:

Class n ruler better ruler worse sum Δword_loss seam_dep base → arm
naht-stark 31 17 14 +0.036888 +7.90° → −0.36°
naht-schwach 29 13 16 −0.000262 +7.44° → −1.18°

The whole ruler loss sits in the class where the eye votes 26 : 2 for the trim. Of the 30 words the ruler punishes, 18 go to the candidate at the eye and none to the base; the two base votes of the entire round fall on words the ruler credits to the candidate. The seam repair is the same size in both classes (+7.90 vs +7.44) — they differ in displacement, not in kink.

Global flip or class rule? Measured, not argued

The narrowing is not an idea but an existing knob (exit_trim_min_kink_deg, the J4b arm). The ladder on the same root:

Threshold words moved naht-stark naht-schwach word_loss seam_dep_median
0° (full J4) 60 31/31 29/29 0.109026 −0.70
54 27/31 27/29 0.109169 −0.10
10° 39 22/31 17/29 0.109460 +2.04
15° 34 21/31 13/29 0.109319 +3.31
20° 19 12/31 7/29 0.109193 +6.52
25° 16 12/31 4/29 0.109134 +7.44
30° 5 4/31 1/29 0.108824 +7.44

No rung separates the classes — it cannot, because their kink is the same. Every rung between 5° and 25° costs MORE than the full trim, and from 10° on the seam repair is handed back. The global flip is the smaller honest mechanism, and trimming the weak class along with it costs nothing: there the ruler is a wash and the eye is 8 : 0 for the trim.

What this PR does and does not claim

adopt: false stands as the tool reported it — the plan needs both thresholds. What the pre-registration licenses for a result >= 60 % is putting it to the author, which is what the entry does. The precedent is the LF11 word round: direction overwhelming, tie threshold failing in every reading, no formal claim, and the author released on that basis. The entry's recommendation is exit_trim as the production default, as a global flip; the switch itself is a rendering-affecting change and therefore the author's call.

Filed like round 6: judgements, narrow key (uid → entry, text, class, repeat_of), analysis JSON and a provenance stamp that also carries the build commit's dirty-worktree flag. The full key, the payload, both arm files and the strata file with its per-word displacements stay out.

menschliche-bewertung.md gains the instrument note as an explicit proposal, not a rule change: a tie threshold defined over all screens also measures the class mix a round chose — and a threshold that follows a round's own class declaration can be softened by cutting the classes differently, which is why it is not adopted here.

Gates

tools.changelog check, tools.docs_register check (1 §14 entry, register agrees), tools.docs_budget check (mess-runde 20 943 of 21 057 — the register grew by its one row and §7.11 was condensed to pay for it, so no budget was raised), pre-commit run --all-files, pytest 2475 passed. Licence side: four text files under an existing, indexed source with a complete SOURCE.md, no binaries, no payloads, tests/test_reserved_history.py green.

🤖 Generated with Claude Code

https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3

Round 5 of the human authenticity pass is judged (2026-09-06, 75 screens,
one sitting). The candidate — the J4 exit trim, `exit_trim` — takes 34 of
the 36 decided screens (94.4 % against a >= 60 % threshold), while the
pre-registered tie threshold of <= 25 % fails at 42.9 %. `adopt: false`
stands as the tool reported it.

The contradiction dissolves per class, and the classes were declared
before the round: `naht-stark` meets BOTH thresholds (26 : 2, 9.7 % ties),
`naht-schwach` runs 8 : 0 for the candidate at 72.4 % ties. That is the
class the pre-registration had described as one where the trim probably
does not show. 21 of the 27 ties come from it; the control class is a
clean 3 of 3.

Re-measured on today's frozen sep05 root, the arm costs +0.000581 on
words, leaves the pairs byte-identical and drops the seam departure from
+7.59 to -0.70 degrees (absolute median 12.67 -> 2.30), doublings 14 = 14.
The whole ruler loss sits in `naht-stark` (+0.036888 against -0.000262),
which is exactly where the eye votes 26 : 2 FOR the trim; of the 30 words
the ruler punishes, 18 go to the candidate and none to the base.

The obvious narrowing was measured, not argued: the existing
`exit_trim_min_kink_deg` knob separates neither class at any rung — both
carry the same kink, +7.90 against +7.44 — costs more than the full trim
between 5 and 25 degrees, and gives the seam repair back. So the global
flip is the smaller honest mechanism, and the entry recommends it; the
switch itself is a rendering-affecting change and therefore the author's
call, as the LF11 round was.

Filed like round 6: judgements, narrow key, analysis JSON and a provenance
stamp that also carries the dirty-worktree flag of the build commit. The
full key, the payload, the arm files and the strata file stay out.

`menschliche-bewertung.md` gains the instrument note as an explicit
PROPOSAL: a tie threshold defined over all screens also measures the class
mix a round chose, and a threshold that follows a round's own class
declaration can be softened by cutting the classes differently — which is
why it is not adopted here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
Copilot AI balanced review requested due to automatic review settings September 6, 2026 21:54

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Provenance and reconstruction claims are inaccurate, and parts of the class analysis contradict their own measurements.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Files and interprets humanbench round 5 for J4 without changing rendering behavior.

Changes:

  • Archives judgments, slim key, analysis, and provenance.
  • Documents class-level results and the tie-gate proposal.
  • Updates J4 tracking and changelog records.
File summaries
File Description
docs/reference/messjournal.md Records round results and analysis.
docs/reference/menschliche-bewertung.md Proposes class-level tie-gate handling.
docs/proposals/tintenfolger.md Updates J4 rescue-path status.
data/humanbench/SOURCE.md Adds round 5 provenance.
data/humanbench/runde-05-vorkommen.json Stores the slim round key.
data/humanbench/runde-05-urteile.txt Stores judgments.
data/humanbench/runde-05-stempel.md Documents build provenance.
data/humanbench/runde-05-auswertung.json Stores aggregate analysis.
data/DATA_PROVENANCE.md Indexes round 5 artifacts.
changelog.d/runde-05-j4-ablage.md Adds the changelog fragment.
Review details

Suppressed comments (4)

data/humanbench/runde-05-stempel.md:41

  • check_arm_scope does not verify this digest. It compares style, source_id, fixture_root, and the arms' exported_at values (tools/humanbench/build.py:1159-1197); wordarm.py does not store a root digest in either arm. Please distinguish the manually recorded digest from the fields the builder actually checked.
| Wurzel-Export | `exported_at 2026-09-02T22:16:06+00:00`, Digest `6cbab9d5c092` (beide Arme, gegeneinander geprüft von `build.py::check_arm_scope`) |

data/humanbench/runde-05-stempel.md:184

  • This deterministic-rebuild claim is not supported while the recorded build has code_dirty: true: the stamp omits the dirty diff, so the seed, root, and commit do not identify the builder code that produced the omitted full key/payload. This should either point to an archived dirty diff or explicitly state that exact reconstruction is unavailable.
Sie bleiben unter `temp/runden-sep04/humanbench/` und sind aus Saat, Wurzel und
diesem Stempel deterministisch wiederherstellbar; die Klassenzuordnung selbst
steht Wort für Wort im schmalen Schlüssel (`stratum`).

data/humanbench/runde-05-stempel.md:199

  • These reconstruction commands omit --fixtures, so they read whichever export currently occupies the default fixture path. The stamp already says the judged round used 6cbab9d5c092 while today's root is eaa195aa7c84; running these commands now therefore creates different arms and a different round. Point both wordarm calls and build at a preserved parent containing the September 2 export, then verify the recorded arm hashes.
uv run python -m tools.humanbench.wordarm --arm "Basis (LF11, Chart-Nib)" \
    --out temp/runden-sep04/humanbench/arm-basis.json
uv run python -m tools.humanbench.wordarm --arm "J4 Austritts-Trim" --exit-trim \
    --registration-from temp/runden-sep04/humanbench/arm-basis.json \
    --out temp/runden-sep04/humanbench/arm-j4.json

data/humanbench/runde-05-stempel.md:120

  • The claim that the dirty tree is harmless contradicts the method's own provenance rule: menschliche-bewertung.md:879-885 says that with code_dirty: true, the commit is only a clue, not proof of the code used. Arm hashes pin geometry, but not dirty builder or strata changes that determine selection, order, and mirroring. Record what was dirty before claiming the commit determines those choices.
> **Der unsaubere Arbeitsbaum gehört genannt, und er ist hier folgenlos.** Was
> die Runde zeigt, sind zwei fertige Arm-DATEIEN mit ihren `sha256`; die Seite
> komponiert nichts nach. Der Bau-Commit bestimmt also die Auswahl, die
> Anordnung und die Spiegelung — nicht die Geometrie. Trotzdem steht das
> Flag hier: ein Stempel, der nur die bequemen Felder trägt, ist keiner.
  • Files reviewed: 10/10 changed files
  • Comments generated: 5
  • Review effort level: Balanced

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread data/humanbench/runde-05-stempel.md Outdated
Comment thread docs/proposals/tintenfolger.md Outdated
Comment thread docs/reference/menschliche-bewertung.md
Comment thread docs/reference/messjournal.md Outdated
Comment thread docs/reference/messjournal.md Outdated
…ot separate

Five findings from the Copilot pass, four of them substantive.

The ladder analysis overstated its own table. `exit_trim_min_kink_deg` DOES
enrich `naht-stark` from 10 degrees on — 71 % of it still firing against 59 %
of the weak class, 39 % against 14 % at 25 — and at 5 degrees the enrichment
even runs the other way. What it never does is SEPARATE the two, which is the
narrower and defensible claim: at 20 degrees it still fires in seven weak words
while already dropping nineteen strong ones, and the reason is the measurement
above it, that both classes carry the same seam kink. The second overstatement
was the closing sentence: 30 degrees does pay 0.000202 of the 0.000581 back, so
the ruler rewards exactly one rung — it just buys that third of the price by
giving up 55 of the 60 words the trim fires on, four strong and one weak
surviving.

The provenance note read as if only rounds 4 and 5 were built on September 4.
All three were; 6 was judged on the 5th, 5 on the 6th, and 4 is still
unjudged. Said that way in the stamp and in SOURCE.md.

The §7.11 row claimed three of the four J4 conversions were done. Two are:
`dspan` and the word round. The arrival side and the Endblenden path are both
open, and the row now says so, which is what §7.9 said all along.

And §8a's Stand block still read "zweimal gefahren" with only the LF11 and J5
rounds in it. It now carries the third and what it did to the tie threshold.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
@MarkusNeusinger

Copy link
Copy Markdown
Owner Author

Review answered in 63a149a — four substantive findings, all of them right.

  • The ladder analysis overstated its own table (messjournal.md line 11283). The threshold DOES enrich naht-stark from 10° on (71 % of it still firing against 59 % of the weak class; 39 % against 14 % at 25°) — and at 5° the enrichment even runs the other way. The defensible claim is the narrower one: it never SEPARATES them. At 20° it still fires in seven weak words while already dropping nineteen strong ones, and the reason sits in the measurement above it — both classes carry the same seam kink (+7.90 against +7.44) and differ in displacement. Rewritten to say enrichment-without-separation, with those numbers in the text.
  • "Nowhere rewarded" was wrong (line 11287). 30° pays 0.000202 of the 0.000581 back, so the ruler rewards exactly one rung. Named as such, together with what that rung costs: it fires on five of 63 words (four strong, one weak) and puts the seam departure back at +7.44°, i.e. it buys a third of the price by giving up 55 of the 60 words the trim fires on.
  • The provenance note read as if only rounds 4 and 5 were built on 2026-09-04. All three were (04/05 at 09:03 UTC, 06 at 10:02); 06 was judged on the 5th, 05 on the 6th, 04 is still unjudged. Corrected in the stamp and in SOURCE.md.
  • §7.11 claimed three of the four J4 conversions were done. Two are — dspan and the word round. The arrival side (1) and the Endblenden path (4) are both open, exactly as §7.9 has them; the row now lists both instead of implying one.
  • §8a's Stand block was stale. It now reads "dreimal gefahren", carries round 5 with its repeat floor taken and its tie threshold broken, and points at the proposal that observation produced.

MarkusNeusinger and others added 4 commits September 7, 2026 00:41
…uffix

The sentence promised a "Nachtrag" while the entry that books it is a
section of its own — a small thing, but the register is read by its shape,
and a pointer that names the wrong shape sends the next reader looking in
this section instead of the one below it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
"Am Lineal teurer" was true of five of the narrowing's six rungs and false
of the sixth — the same overreach the review caught two paragraphs above,
one sentence later. Say five of six, and say what the sixth does instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3
# Conflicts:
#	docs/reference/messjournal.md
# Conflicts:
#	docs/reference/messjournal.md
@MarkusNeusinger
MarkusNeusinger merged commit 61fed8f into main Sep 6, 2026
8 checks passed
@MarkusNeusinger
MarkusNeusinger deleted the runde-05-j4-ablage branch September 6, 2026 23:21
MarkusNeusinger added a commit that referenced this pull request Sep 6, 2026
**The owner approved this on 2026-09-06 ("1 ja") — booked as author
decision A37.** Stacked on #554, which files the round this decision
rests on; **merge #554 first**, this PR's base is its branch.

Turns `exit_trim` into the production default. The switch stays — the
pre-adoption base is a bench arm like any other — but production now
writes on the other side of it, and every `/write/word` answer moves
with it (the edge cache holds the old one for up to 24 h, as after the
LF11 write). No DB row is touched: this decision flips a rule, it does
not write geometry.

## The re-baseline, on an UNCHANGED root

`--expect-root eaa195aa7c84,0fbde2d72b64`, `OPENBLAS_NUM_THREADS=1
OMP_NUM_THREADS=1`. The root does not move, so these two lines are
paired, not merely consecutive:

| | words | pairs | `seam_dep_median` | abs. median | `seam_arr_median`
| `gleichzug_doublings` |
|---|---|---|---|---|---|---|
| before (`--no-exit-trim`) | 0.108444 | 0.148236 | +7.59° | 12.67 |
−2.83° | 14 |
| **after (default)** | **0.109026** | **0.148236** | **−0.70°** |
**2.30** | −3.87° | 14 |
| Δ | **+0.000581** | **0.000000** | −8.29° | −10.37 | −1.04° | 0 |

63/63 words and 33/33 pair drills scored, none skipped, none failed.
`worst_word` moves `han` 0.232609 → `regieren` 0.233052; `worst_pair`
stays `In` 0.283819. Components: `comp_transition` 0.089804 → 0.091028,
`comp_coverage` 0.101580 → 0.101667, `comp_width` 0.162397 unchanged —
the trim touches the seam, not the width. The doublings at the delivered
nib stay at 14, which was not a given for a rule that takes ink away.

**The word ruler rises knowingly.** Round 5 showed the whole of that
cost sits in `naht-stark`, the class where the eye votes 26 : 2 **for**
the trim; of the 30 words the ruler punishes, 18 go to the candidate at
the eye and none to the base. `EXIT_TRIM_MIN_KINK_DEG` stays 0.0 — the
same round measured the narrowing and it separates neither class at any
rung.

## The golden fixture: a declared re-baseline

Re-baked with `REGEN_GOLDEN=1`. Measured before the regen:

- **10 of the 11 words move**, `wovon` does not (its exits are backward,
bow and arm exits, which the class excludes by construction).
- **No word gains or loses a draw item** — item counts 6 … 18 before and
after; 2 to 8 items move per word, namely the trimmed letters and their
connectors.
- **A solitary glyph with `pen=None` stays byte-identical** — checked
across all 23 glyphs in the golden payloads, 0 move, because the rule
cannot fire without a following connector. The CLAUDE.md invariant holds
as written.

## Two measurement layers had to follow, and one is a repair

**`pairlab.prodconn` had named this exact case itself.** Its docstring
said the trim replaces the generator's return value inside
`compose_word`, "harmless while the switch is off — should it ever
become the default, this function has to grow the same post-processing
or the dissection will quietly measure the wrong curve." It now does.
The recorder additionally captures `_cut_exit_stub`, which runs if and
only if the trim was really applied (guards, collinear cut and the
min-kink narrowing all sit in front of it), so what is recorded is the
DECISION rather than a re-derivation of it; `replay` shifts that stub
with its letter and calls `_exit_trim_index`/`_straight_to` again. The
rule is never restated in `tools/`. The parity tests now say it in two
halves: an untrimmed join reproduces the recorded call point for point,
a trimmed one reproduces the rule's own two invariants — the coupling
point does not move, and the join is a straight line.

**The chain is exempt, and that is measured rather than assumed.** Its
init default is the frozen mirror (the production init was measured as
K-F on `sep04` and rejected), and its guard reads the composition soll.
`ductus_soll` over all 63 word samples returns **126 rows of which 0
move**: crossings 292, zones 177, strokes 111, touches 89, overlaps 11
on both sides. The trim cuts a stub and straightens a connector; it does
not touch the composed word's topology, so `k0eval`, the structure guard
and every chain number stand on unchanged ground.

**What is left due, named rather than done here:** the pilot's map IS
the composed path, so its `sep05` dev-19 numbers now stand on a
composition that no longer exists. That is a route measurement with its
own protocol (`/verify-trace`), not something an adoption PR should fold
in — it is filed as an open arm in `tintenfolger.md` §7.11 and is due
before the next pilot statement.

## Docs and gates

`messjournal.md` gains the dated adoption entry plus its register row
and a headline-ledger row; `qualitaetsmetrik.md`'s status block carries
the new headline (and lost two lines to stay under its own 40-line cap);
the glossary entry for Austritts-Trim flips from "opt-in, not adopted"
to adopted with both instruments' numbers; `werkzeuge.md` and
`write-api.md` follow.

`mess-runde` is raised to 23 567 — measured 21 425 plus the documented
10 %. The reason is the one the budget's own comment already licenses:
the register books the standard PAIR of rows for a round and the
adoption it triggered. Both rows were condensed to roughly 300 tokens
each (they stood near 500, three times the register's ~163-token
average) and the §7.11 J4 row was rewritten SHORTER than it was, so the
residual prose growth is one new open-arm row and 11 tokens in a Stand
block.

Gates green locally: `pytest` 2478 passed / 8 skipped, `pre-commit run
--all-files`, `tools.changelog check`, `tools.docs_register check` (2
§14 entries, register agrees), `tools.docs_budget check`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01UEScQMZFvxxNNyNJYryfa3

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants