docs(m4): start-of-stage brief — vocabulary collateral strip, D19–D22 for freeze - #10
Conversation
… for freeze M4 added by Kyle's post-M3 pick (2026-07-28): close M3's stated bound (collateral among 12, not across the vocabulary) before write-up + /seed-hunt. S1/S2 stretches declined and banked as idea #13 in the j-lens-proj-ideas backlog. Brief contents: 12 primes x 180 probes frame (D19), single-clause VOCAB-SPARING level gate on the never-measured non-subset pool (D20), treatment of the four cross-mention confound cells found while drafting (D21), D18 bars widened to all 60 scored words (D22). Realized ns computed from recorded artifacts with the frozen oracle; both inherited obligations dispositioned explicitly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANtkrPFv6i9CgCasojqcoZ
F1: own the 0.5 gate constant as new, uncalibrated and deliberately lenient rather than "carried from M3" — the statistic changed (12-fold conjunction vs per-cell; 0.5 on the conjunction ≈ 0.944 per-cell under independence) and the status changed (D17 qualifier that could never rescue a claim → sole dispositive gate). The per-cell equivalence and the deletion-count dependency go into GATE_WORDING itself; new deviations row. F2: own D9(b)'s span-truncation residual, which re-enters a gate-bearing arm for the first time since M1 now that probes widen to all 60 — 0 / 3 / 4 span-filling gate-arm items (1.5B trumpet-1, beetle-1, butterfly-1; 3B trumpet-1/2/3, butterfly-1), biasing toward the gate where the 1.5B margin is one item. Adds an instrument fact, a power-table row, a deviations row and a pre-registered residual-conservative recomputation (never dispositive), and corrects the claim that D22 pins the oracle's premises — it pins the span premise, not the scope premise. oracle.py untouched. F3: restate zero-gating (25 / 7 / 5 concepts) as competence-selection, not tokenizer geometry — under D9(b) the gate is token-count-insensitive up to 3 and the recorded clean spans are wrong answers, not truncations. States the upward bias on the floor and connects it to F12's selection-enrichment. F4: correct "cells no milestone has ever measured" to 486 / 844 / 993 (of 492 / 852 / 1,008); the remaining 6 / 8 / 15 are M1-recorded — the very cells used as the proxy — so they are in-sample and cap the gate arm at 35/41, 69/71, 82/84 before any forward pass. Power-table row split, ceilings stated as pre-registered fact, out-of-sample framing tempered. HANDOFF.md mirrors the F1 and F4 corrections. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Both are defects of the F2 fix, not of the finding it addressed.
F11: the residual-conservative selector was degenerate and its counts were
wrong. "The span fills the 3-token window" matches every recorded cell —
greedy_3 IS the 3-token span. The real condition is narrower: the span, after
stripping leading whitespace, equals the concept's spelling with nothing
following it, so no boundary is observed. On that reading the gate-arm
residual is 0 / 2 / 2, not 0 / 3 / 4 — 1.5B beetle-1, butterfly-1; 3B
trumpet-3, butterfly-1 — which is exactly the set oracle.py's frozen docstring
already names ("beetle-1, butterfly-1 x2 arms, trumpet-3"). trumpet-1 (1.5B)
and trumpet-1/-2 (3B) record 'Trumpet<|im_end|>' and <|im_end|> closes a word
under oracle._BOUNDARY, so they carry no residual; generated 'Trumpet' is two
tokens where bare 'trumpet' is three, which leaves room for the terminator.
Corrected in all three places (instrument fact, power table, deviations row).
F12: the conservative read's denominator was left to the runner, and the two
readings straddle the bar (same numerator: wilson(43,71) lower 0.489 fails,
wilson(43,69) lower 0.505 passes). Pre-registers fail-in-place — the arm stays
at its as-scored n (41 / 71 / 84, preserving the power table's cross-check)
and residual-affected items score as failures — and names and rejects the
un-gating alternative as the less conservative of the two.
F13-F15 (nice-to-have) remain FOLLOW-UP pending Kyle's pull-in call.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Kyle's call at the freeze gate: pull all eight in before D19-D22 freeze,
since four of them get byte-frozen into code once he signs.
F5: re-scan the cross-mention confound with D5's OWN rule (prefix +
forbidden_forms, not whole-word) — five pairs, not four. Adds Egypt->beetle-2
("Ancient Egyptians carved amulets of the scarab..."), which a whole-word scan
misses because the clue inflects the prime. beetle-2 is ungated on all three
subjects ('Scarab'/'scarab'), so it carries no gate-bearing cell today; it is
listed so a future re-gate cannot silently acquire one. D21 retitled, list
frozen in its text, deviations row and power-table row updated.
F6: replace the proxy-only 0.5B prediction with the in-statistic number.
M3's recorded 0.5B subset reads 19/28 = 0.679 under M4's own gate statistic,
Wilson lower 0.4934 — it ALREADY FAILS M4's bar. So there is no "tension"
between M3's clean 0.5B floor and an M4 failure: M3's floor was clean only
under the cluster-mean per-cell statistic, and switching to the conjunction
flips it. Sharpest illustration of D20's point that the constant carries no
meaning across statistics.
F7: move the realized ns outside the underpowered conditional — the proposed
wording literally read "41 / 71 / 84 < 20". Now abstract, with realized n in a
trailing parenthetical, matching m1_battery.GATE_WORDING's shape.
F9: add the re-certification precondition to GATE_WORDING. The single-clause
level bar, read alone, is satisfiable by a no-op intervention (~100% survival
prints VOCAB-SPARING). The guarantee existed in D19's design and the runner's
exit code but not in the sentence a write-up quotes; now the sentence carries
its own precondition.
F10: PROJECT.md Next action and README status now point at M4 pending freeze
instead of contradicting HANDOFF in the same tree.
F13: name all three zero-gating mechanisms, not one — answers something else,
answers correctly behind a modifier the opening-word rule refuses ('Peking
duck', 'Electric guitar', 'Insect Beetle'), or misses on morphology. The
upward-bias conclusion is unchanged.
F14: the 0.5B ceiling (35/41, Wilson lower 0.716) is far above the bar and is
not a reason to expect failure; the reasons are M3's in-statistic 19/28, the
0/6 proxy and M1's 33/69 control cell. Removes a "so" that contradicted the
preceding sentence.
F15: re-wrap the paragraphs the earlier fix commits broke.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
F16 [should-fix]: the F11 fix swapped a selector matching EVERY cell for one matching NONE. "Equals the scored concept's spelling exactly" is case-exact; recorded residual spans are 'Beetle' / 'Butterfly' / 'Trumpet' while roster spellings are lowercase, so a verbatim runner would select zero cells and turn the sole mitigation for a bias-toward-the-gate into a no-op. Verified: over all 540 cells x 3 subjects, case-exact selects 0; case-insensitive selects exactly the pre-registered 0 / 2 / 2. The selector now states case-insensitive comparison "exactly as oracle.says_concept_prefix compares (re.IGNORECASE)", and says why the case rule decides the set rather than leaving it as a detail. F17: the F10 fix rewrote each file's forward-looking prose but not the status lines above it, leaving PROJECT.md and README.md self-contradictory (each was internally consistent before). PROJECT.md's Status paragraph and README's Status line now both carry M4-in-flight, and README's milestone table gains an M4 row with S1/S2 relabelled banked. F18: "exactly the set oracle.py's docstring names" was four cells against the oracle's stated six — the docstring enumerates across arms (butterfly-1 x2 arms: clean and control_late). Restated as the gate-arm members of that set, with the arm/subject scoping spelled out. F15 re-fix: the :383 wrap was missed and the F13 fix introduced a new orphan at :105. Both repaired; full re-scan of all four files is clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Kyle froze all four decisions 2026-07-29 at the review gate: D19(a) full 12 x 180 strip + clean re-run (2,340 cells/subject), with the embedded 255-cell M1 and 468-cell M3 re-certifications graded first. D20(a) the single-clause survives-all-12 level gate as written. D21(a) all five cross-mention pairs kept in the gate-bearing pool with named per-cell reporting. D22(a) both run-time instrument bars over all 60 scored words + the 12 direction words. F8 (the verdict-label choice, deliberately left open as a design question rather than fixed as a defect) resolved: the label stays VOCAB-SPARING and the realized survival proportion rides INSIDE the verdict string — "VOCAB-SPARING (k/n survive all 12 = <rate>; Wilson 95% lower <lo>)". This is M3's ON A DAMAGED FLOOR pattern applied to a level bar: the label is what a write-up quotes and it is frozen into every results JSON, so the number that qualifies it travels with it rather than living in prose. The brief now states why — at the 1.5B pass point 27 of 71 gated items (38%) are damaged by at least one deletion, and at the nominal 0.5 it would be half. Freeze note added to the brief per the M3 precedent; full DECISIONS.md entries D19-D22 land with the M4 code PR, per the M0-M3 pattern. PROJECT.md next action moves to building the runner; HANDOFF.md records the freeze. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Line-by-line patching kept shifting the break point; re-flowed both whole paragraphs at 78 columns instead. Also drops a line in HANDOFF that still said "code only after freeze" now that D19-D22 are frozen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
F19: the freeze commit updated PROJECT.md and HANDOFF.md but not README.md, so
the front page still said "decisions not yet frozen" while the other two said
FROZEN — in a repo whose whole discipline is freeze-before-code, that answered
"were D19-D22 frozen before the run?" with "no". Both README spots corrected.
F20 [should-fix, amends the frozen D20 wording]: the verdict-string clause as
frozen carried only the AS-SCORED proportion, so the two pre-registered reads
that can flip which number is honest — residual-conservative fail-in-place and
the concept-level collapse, each closed in the brief with "the ... numbers are
the honest ones to quote" — stayed in prose. That is precisely the failure the
F8 resolution was adopted to prevent ("prose is not what gets quoted; the label
is"). The window is live: at 1.5B the bar needs k >= 44 (wilson(44,71) lower
0.50342), and fail-in-place removes the 2 residual items, so at an as-scored k
of 44 or 45 the JSON would have printed VOCAB-SPARING ... lower 0.503 while the
brief's own pre-commitment said 42/71 (lower 0.47539) was honest. M3's
precedent is a CONDITIONAL qualifier attached by the runner (m3_matrix.py:179);
M4 had borrowed the sentence shape and dropped the mechanism. D20 now carries a
pre-declared AS-SCORED ONLY qualifier, attached whenever a conservative read's
Wilson lower bound is below 0.5 while the as-scored read's is not, naming each
such read and its number. Per D17's carried rule the qualifier scopes a claim
and can never create or rescue one; the gate, its 0.5 bar, its arm and its
re-certification precondition are all unchanged.
F21: "the verdict, whatever it is" was followed by a template hard-coding the
pass label, and no failing label appeared anywhere in the brief. NOT
VOCAB-SPARING is now named in the wording.
Recorded in the brief's freeze note as an amendment POST-FREEZE, PRE-RUN — the
M3 precedent (its own round-4 F15/F16 amendment) — and FLAGGED FOR KYLE'S
RATIFICATION IN THE MERGE BRIEF.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Kyle ratified the AS-SCORED ONLY qualifier amendment 2026-07-29 ("I ratify the
F20 amendment"), so the freeze note now records it as ratified rather than
flagged-pending. The amendment itself is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
…e same clause F22 [should-fix]: the F21 fix invented NOT VOCAB-SPARING as the failing label. All three predecessors emit the lineage's pre-committed null "not shown" (m1_battery.py:662, m2_depth.py:857, m3_matrix.py:1124), and for good reason — failing a Wilson LOWER bound does not establish the contrary, so an assertive negative cannot be earned by it (at 1.5B, k=40 has a point estimate of 0.563 above the bar with a straddling interval, yet would have printed the negative). The wording now emits "not shown", and scopes the string: the gate verdict is the AND over the two gate-bearing subjects, while 0.5B is reported in the same shape under its standing frame and is never a gate claim — the brief had pre-committed the string "per subject" while the same block says 0.5B is never gate-bearing and the power section expects 0.5B to fail. F23 [should-fix]: HANDOFF.md never got the round-4 amendment — still "three adversarial rounds", still summarizing the pre-amendment string F20 found defective, still reading round 4 as pending, and still "D21 decides their treatment" after D21 froze. Fourth recurrence of the F10/F17/F19 propagation class; the whole M4 paragraph is rewritten and the file's Last-updated line moved to 2026-07-29. Three nice-to-haves are fixed here rather than deferred BECAUSE THEY LIVE IN THE SENTENCES F22 FORCED ME TO REWRITE — leaving known defects in a clause I was rewriting anyway would be perverse, and all three are in the frozen wording that byte-freezes into code. Flagged for Kyle rather than assumed: - F26: "the verdict, whatever it is, additionally carries the qualifier" had no precedence exclusion, so verbatim it attached AS-SCORED ONLY to DEGENERATE / UNDERPOWERED / NOT A RESULT — contradicting the D17 rule the same sentence cites. Now explicitly bar-level only, matching m3_matrix.py:1109-1125, and the D20 precedence paragraph says so too. - F25: the frozen example had dropped "survive all 12" from the template, the qualifier had an example but no template, and the both-reads-fire rendering was unspecified — against the brief's own "stated once and implemented verbatim" rule. Both templates now stated explicitly, with a fixed read order. - F24: the "k of 44 or 45" window quoted only k=44's numbers. Both are now worked (44 -> 0.503 vs 42/71 0.475; 45 -> 0.518 vs 43/71 0.489). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
F27 [should-fix]: the F22 label swap reached every use of NOT VOCAB-SPARING except the amendment note that documents it, leaving the retired label as the only surviving assertion of itself in the repo — inside the paragraph the DECISIONS.md D20 entry gets built from. Fifth recurrence of the F10/F17/F19/F23 partial-propagation class. F28 [should-fix]: a427b2d was a SECOND material amendment to the byte-frozen GATE_WORDING block, but the freeze note recorded only the round-4 amendment and carried only its ratification quote ("I ratify the F20 amendment"). The unratified block therefore read as covered by Kyle's ratification. It is now recorded as a separate, explicitly NOT-YET-RATIFIED Amendment 2, enumerating what changed against eef9a27: (i) the failing label became the lineage null "not shown"; (ii) 0.5B is scoped inside the wording; (iii) the qualifier attaches to claim-level verdicts only; (iv) both string templates are stated. The claim that the gate, its 0.5 bar, its arm and its precondition are unchanged is re-verified byte-for-byte against 90b994c and holds under both amendments. HANDOFF carries the same split. F29: HANDOFF's "reviewed across five rounds" was true only until this round's triage; replaced with a pointer to the mailbox, which cannot go stale. Amendment 2 is flagged for Kyle's ratification; nothing merges until he rules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Kyle ratified the second post-freeze amendment 2026-07-29 on the recommendation that enumerated all four of its changes: the failing label became the lineage null "not shown"; 0.5B was scoped inside the wording; the AS-SCORED ONLY qualifier attaches to claim-level verdicts only; both string templates are stated explicitly. The amendment itself is unchanged — only its status. With this, both amendments to the frozen D20 wording package are ratified, and the gate, its 0.5 bar, its arm and its re-certification precondition stand exactly as frozen at 90b994c. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL
Adversarial review — CLEAR TO MERGESix reviewer rounds over Rounds 4, 5 and 6 were Kyle-authorized overages beyond the three-dispatch cap, each granted because the preceding round had found a defect inside the previous round's fix. Disposition
Waivers, verbatim
— on a recommendation that (1) ratified Also recorded: "I choose option b. I authorize a fourth reviewer dispatch", "I agree with your recommendations. (a) across the board", "I ratify the F20 amendment and approve a fifth round", "one more round like you mentioned". Decisions frozen in this PRD19–D22, (a) across the board (Kyle, 2026-07-29). Two post-freeze, pre-run amendments to the D20 wording package, both ratified; the gate, its 0.5 bar, its arm and its re-certification precondition are unchanged from the freeze at The pattern worth carrying into the code PREvery defect after round 1 was the same shape — a rule stated in prose that breaks when read literally by a runner: F11's selector matched everything, F16's matched nothing, F20's string carried the number the brief calls dishonest, F22's label asserted what a lower bound can't establish. Follow-ups carried: none. Every finding is verified, waived, or closed. |
What this is
The start-of-stage brief for M4 — the vocabulary collateral strip, the close-out stage Kyle picked on 2026-07-28 (post-M3 decision session) over closing immediately and over the S1/S2 stretches (declined and banked as idea #13 in the j-lens-proj-ideas backlog). Docs only — no runner code until Kyle freezes D19–D22.
What M4 is
M3's Honest limits states the bound: the matrix measures collateral among 12 concepts, not across the vocabulary. M4 closes it — the 12 characterized primes run against all 180 M1 items; the new content is the never-measured non-subset pool (48 concepts' gated items × 12 deletions: 492 / 852 / 1,008 cells per subject).
Decisions proposed for freeze
Both inherited obligations dispositioned explicitly:
oracle._BOUNDARYnot triggered (no non-ASCII vocabulary), and the conjunction-degeneracy carry-forward discharged by a deliberately single-clause gate whose one surviving arm is named in the wording.All ns realized from recorded artifacts with the frozen oracle (gate arm 41 / 71 / 84; worked bar at 1.5B: ≥ 44 of 71, computed with the frozen ruler). HANDOFF updated to the new state.
🤖 Generated with Claude Code
https://claude.ai/code/session_01ANtkrPFv6i9CgCasojqcoZ