diff --git a/HANDOFF.md b/HANDOFF.md index 50c183d..45bc701 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -1,6 +1,6 @@ # HANDOFF.md — mute-map -_Last updated: 2026-07-28_ +_Last updated: 2026-07-29 (M4 brief reviewed; D19–D22 frozen)_ ## What was just done @@ -78,24 +78,52 @@ and each needs its own brief. ## Immediate next move -**Nothing is pre-committed.** The honest options, none of them owed: - -1. **Close the project** — a write-up of the four-milestone arc, then - `/seed-hunt` for the next paper. The characterization KICKOFF bought is - delivered: breadth (M1), localization + dose (M2), specificity (M3). -2. **S1 stretch — scale.** 7B lens fit on a rented GPU (≤ $15, decision K3's - no-refit rule applies only to the core chain) plus a matrix-lite. The - sharpening-with-scale story is the one M1–M3 keep gesturing at and never - measured above 3B. -3. **S2 stretch — scope.** Token mute button or concept mute button? - (translations, synonyms, morphological variants). **Its brief owes the - `oracle._BOUNDARY` boundary-class decision before it freezes any non-ASCII - list** — that is the named future trigger D18 recorded, and it is the only - inherited obligation on the board. -4. **The cheap follow-up M3 explicitly did not run:** M3 measures collateral - among 12 concepts, not across the vocabulary. Nothing here shows that deleting - France spares the other 48 M1 concepts. A 12-prime × 60-probe strip would - close that gap for roughly the cost of one M3 subject run. +**Decided (Kyle, 2026-07-28, session "mute-map post M3 decisions"): run the +vocabulary collateral strip as close-out stage M4, then write-up + +`/seed-hunt`.** The S1/S2 stretches were declined for this repo and banked as +idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md` — they +compete in the seed-hunt on equal terms, no incumbent's privilege. + +**Where M4 stands: `docs/M4-BRIEF.md` is written, adversarially reviewed (see +the mailbox under `~/.claude/reviews/mute-map/` for the round-by-round +record), and D19–D22 are FROZEN (Kyle, 2026-07-29) — (a) across the board.** +The verdict label was resolved per review F8: the label stays `VOCAB-SPARING` +and the realized survival proportion rides inside the verdict string, the M3 +`ON A DAMAGED FLOOR` pattern applied to a level bar. **Amended post-freeze, +pre-run and ratified by Kyle (round 4, F20):** that string as first frozen +carried only the as-scored proportion, so the two pre-registered reads that +can flip which number is honest (residual-conservative fail-in-place; +concept-level collapse) stayed in prose — the exact failure F8 was adopted to +prevent. D20 now carries a pre-declared **AS-SCORED ONLY** qualifier, attached +by the runner whenever a conservative read's Wilson lower bound falls below +0.5 while the as-scored read's does not. **Amendment 2 (rounds 5–6, +F22/F24–F27), also ratified by Kyle 2026-07-29:** the failing label became the +lineage's null `not shown` rather than an assertive negative, 0.5B was scoped +inside the wording, the qualifier's attachment was restricted to claim-level +verdicts, and both string templates were stated explicitly. The gate, its 0.5 +bar, its arm and its re-certification precondition are unchanged under both +amendments. Full `DECISIONS.md` entries D19–D22 land with the M4 code PR, per +the M0–M3 pattern. + +The stage: 12 subset primes × all 180 M1 items (2,340 cells/subject), gate = a +single-clause VOCAB-SPARING level bar on the non-subset pool (per-item +survives-all-12, Wilson lower bound ≥ 0.5 at 1.5B AND 3B — a *new*, +deliberately lenient, uncalibrated constant on the 12-fold conjunction, ≈ +0.944 per-cell under independence, explicitly not carried from M3's per-cell +floor). Realized ns are known from the recorded gated sets (gate arm 41 / 71 / +84); 486 / 844 / 993 of the pool's 492 / 852 / 1,008 cells are genuinely new, +the rest are M1-recorded and cap the gate arm at 35/41, 69/71, 82/84 before +any forward pass. Design facts found while drafting: five probe clues mention +a prime's spelling under D5's own prefix rule (October→september-2, +silver→flute-1, China→jade-1, October→opal-2, Egypt→beetle-2 — the last +ungated on all three subjects; D21(a) keeps all five in the pool); D9(b)'s owned +span-truncation residual re-enters a gate-bearing arm for the first time since +M1 (0 / 2 / 2 cells, carried by a pre-registered fail-in-place conservative +read); all 60 roster words pass the D18 span bar on all three tokenizers +(checked in advance); and M3's own 0.5B subset already fails M4's bar +in-statistic (19/28, Wilson lower 0.4934), so the off-gate 0.5B floor may +genuinely fail — reportable under the standing frame. No runner code exists +yet; with D19–D22 now frozen, code is the next step, cut from `m3_matrix.py`. Standing constraints unchanged: certified environment = `mps` + torch 2.13.0 + transformers 5.13.1 (off it: NOT A RESULT); `m0_anchor.py` stays certified and diff --git a/PROJECT.md b/PROJECT.md index fd9fead..9137215 100644 --- a/PROJECT.md +++ b/PROJECT.md @@ -9,13 +9,20 @@ deleting one concept's direction at the late third silences that concept and leaves the other eleven almost untouched (1.5B: diagonal 0/34 vs off-diagonal 363/374). M2 PASSED (LATE-LOCALIZED at 1.5B and 3B), M1 PASSED (BREADTH-SPECIFIC at 1.5B and 3B), M0 PASSED (2026-07-27); all re-certified -bit-for-bit on every later run. v1 = M0–M3 per `docs/KICKOFF.md`; S1/S2 -stretches optional. +bit-for-bit on every later run. v1 = M0–M3 per `docs/KICKOFF.md`. **Close-out +stage M4 (the vocabulary collateral strip) is now in flight** — brief written +and reviewed, **decisions D19–D22 frozen 2026-07-29**, no runner code yet. The +S1/S2 stretches were declined for this repo and banked (idea #13). -**Next action:** none pre-committed — the v1 chain is closed. Open options: the -S1 stretch (7B lens fit + matrix-lite, ≤ $15 rented GPU), the S2 stretch -(lexical vs semantic scope, which owes the `oracle._BOUNDARY` boundary-class -decision before freezing any non-ASCII list), a write-up, or `/seed-hunt`. +**Next action:** build the M4 runner. `docs/M4-BRIEF.md` is written and +adversarially reviewed (PR #10); **D19–D22 were frozen 2026-07-29 — (a) across +the board** — so the next step is code: `m4_strip.py` cut from `m3_matrix.py`, +`m4_verdict.py`, `test_m4.py`, and the D19–D22 entries appended to +`docs/DECISIONS.md` in that same code PR. After M4: write-up + `/seed-hunt`. +The S1 (7B) and S2 (lexical vs semantic scope) stretches were declined for +this repo and banked as idea #13 in +`~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md`; they compete in the +seed-hunt on equal terms. **Key facts** - Fact — Anchor: S4b (dim-stage), concept-specific off-switch at 1.5B, +.727 diff --git a/README.md b/README.md index 818e4c3..aa8638e 100644 --- a/README.md +++ b/README.md @@ -21,27 +21,30 @@ pre-registered gates frozen as code before any run): | M1 | Breadth | How much of the (measurable) vocabulary has an off-switch? | | M2 | Localization + dose | Where does the switch live, and how much removal does it take? | | M3 | Specificity | The full prime × probe collateral matrix. | -| S1 (stretch) | Scale | The specificity-emergence curve, extended to 7B. | -| S2 (stretch) | Scope | Token mute button or concept mute button? | +| M4 (close-out) | Vocabulary collateral | Does deleting one concept spare the *other 48*? | +| S1 (banked) | Scale | The specificity-emergence curve, extended to 7B. | +| S2 (banked) | Scope | Token mute button or concept mute button? | **The honest framing:** an effect *found during a replication, characterized here* — the anchor is dim-stage's own recorded result, not a paper claim. -**Status: M3 PASSED 2026-07-28 — the v1 chain is complete.** On the full 12 × 12 -prime × probe matrix at the switch's home band, deleting one concept's direction -silences that concept and leaves the other eleven almost untouched: at 1.5B the -diagonal names **0/34** while the pooled off-diagonal names **363/374** -(+0.971 [+0.867, +0.983]); at 3B 3/32 vs 343/352 (+0.881 [+0.731, +0.943]). The -same contrast restricted to *same-category* pairs — the arm the lineage's single -control actually tested — is +0.950 and +0.891, likewise CI-clean. Of the -matrix's 132 ordered off-diagonal pairs, **126 had never been measured before** -(the other 6 are M1's own country control cells, which this run re-certifies -bit-for-bit before it reads anything new). What the gate did not -ask: collateral concentrates on a few fragile **probes** rather than being caused -by damaging **primes** (`silver`'s direction damages *nothing*, while `silver` -itself is the most fragile probe in the grid — inverting what a single control -cell had suggested), category-block collateral is CI-clean at 0.5B and dissolves -by 1.5B, and the only imperfect mutes anywhere are `Egypt` and `October` at 3B — +**Status: M3 PASSED 2026-07-28 — the v1 chain is complete; close-out stage M4 +is in flight — brief reviewed, decisions D19–D22 frozen 2026-07-29, runner not +yet written.** On the full 12 × 12 prime × probe matrix at the switch's home +band, deleting one concept's direction silences that concept and leaves the +other eleven almost untouched: at 1.5B the diagonal names **0/34** while the +pooled off-diagonal names **363/374** (+0.971 [+0.867, +0.983]); at 3B 3/32 vs +343/352 (+0.881 [+0.731, +0.943]). The same contrast restricted to +*same-category* pairs — the arm the lineage's single control actually tested — +is +0.950 and +0.891, likewise CI-clean. Of the matrix's 132 ordered +off-diagonal pairs, **126 had never been measured before** (the other 6 are +M1's own country control cells, which this run re-certifies bit-for-bit before +it reads anything new). What the gate did not ask: collateral concentrates on +a few fragile **probes** rather than being caused by damaging **primes** +(`silver`'s direction damages *nothing*, while `silver` itself is the most +fragile probe in the grid — inverting what a single control cell had +suggested), category-block collateral is CI-clean at 0.5B and dissolves by +1.5B, and the only imperfect mutes anywhere are `Egypt` and `October` at 3B — the two concepts pre-registered as the leaky-switch stratum. **M2 PASSED 2026-07-28** — on a pre-registered 12-concept subset, the @@ -70,9 +73,15 @@ deterministic prefix rule on the recorded 3-token span (decision D9b, frozen before any run), M1's published numbers stand untouched, and the re-score is published beside them as a labelled reanalysis (D10a) in which the contrast survives on every subject and the dark categories light up. **The v1 chain -(M0–M3) is now closed; the S1 (7B) and S2 (lexical vs semantic scope) stretches -remain optional.** Models: Qwen2.5-0.5B/1.5B/3B-Instruct, local MPS, -forward-only; core chain $0. +(M0–M3) is now closed.** In progress: close-out stage **M4, the vocabulary +collateral strip** (12 characterized directions × all 60 concepts), which +measures the one thing M3's near-white grid does *not* show — that deleting +France spares the other 48 concepts. Its brief is adversarially reviewed and +its decisions (D19–D22) were frozen 2026-07-29, before any runner code exists +— the lineage's freeze-before-code discipline. The S1 (7B) and S2 (lexical vs +semantic scope) stretches were declined for this repo and banked for a future +seed-hunt. Models: Qwen2.5-0.5B/1.5B/3B-Instruct, local MPS, forward-only; +core chain $0. Full brief: [`docs/KICKOFF.md`](docs/KICKOFF.md). The 12-idea backlog this was picked from: dim-stage diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md new file mode 100644 index 0000000..6ac6e67 --- /dev/null +++ b/docs/M4-BRIEF.md @@ -0,0 +1,647 @@ +# M4 start-of-stage brief — the vocabulary collateral strip + +*Start-of-stage brief per the per-stage rhythm: plain-terms explanation first, +design extraction second, decisions third, code only after Kyle freezes. +Decisions here are D19–D22, continuing `docs/DECISIONS.md` (D15–D18 were M3's).* + +*Provenance, stated up front: M4 is **not** in KICKOFF's frozen chain. The v1 +chain (M0–M3) closed on 2026-07-28; this stage is the close-out follow-up +HANDOFF listed as option 4, picked by Kyle the same day (session "mute-map post +M3 decisions") over closing immediately and over the S1/S2 stretches — which +were banked as idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md`. +KICKOFF's scope decisions are not relitigated; this is an addition, owned in the +deviations table. After M4, the plan of record is write-up + `/seed-hunt`.* + +## What M4 is, in plain terms + +M3's killer figure shows a dark diagonal on a nearly-white grid: deleting any +of the 12 subset concepts' directions mutes that concept and almost nothing +else — *among those 12*. M3's own Honest limits section states the bound +plainly: nothing measured so far shows that deleting France spares the other +48 M1 concepts. The natural over-reading of the figure ("the deletion spares +the vocabulary") is exactly one experiment wider than what was run. + +M4 runs that experiment. Keep the 12 characterized directions as the **primes** +(the thing deleted); widen the **probes** (the thing asked about) from the 12 +subset concepts to **all 60 M1 concepts** — the full frozen 180-item battery. +Every cell is the M3 recipe unchanged: ablate prime A's direction at the +subject's late third (λ = 1, k = 1), ask item-of-B its naming question, score +with the frozen D9(b) oracle. The strip is 12 × 180 = 2,160 ablated cells per +subject; the genuinely new content is the **non-subset pool** — the 48 concepts' +gated items under each of the 12 deletions: **492 / 852 / 1,008** cells at +0.5B / 1.5B / 3B, of which **486 / 844 / 993** have never been measured by any +milestone. The remainder (6 / 8 / 15) are already recorded in M1 — they are the +non-subset items whose frozen M1 control direction happens to be one of the 12 +primes — so those outcomes are fixed before the run and are stated below as +pre-registered ceilings, not predictions. + +The strip is also a mostly out-of-sample test of M3's most interesting +descriptive finding. Finding 1 said collateral concentrates on a few fragile +**probes** (silver's column) rather than on damaging **primes**. The 48 new +probe columns are a near-fresh sample — 6 / 8 / 15 of their gate-bearing cells +are M1-recorded, the other 486 / 844 / 993 are new: either near-total sparing +holds and the fragile-probe story stays a curiosity of the 12, or more +silver-like columns exist among the 48 — both outcomes are findings, and a +*failed* gate (a low survival floor across the wider vocabulary) would be the +pre-committed reportable headline against the specificity story's reach. + +Because the strip contains, cell-for-cell, almost everything M1 and M3 already +recorded on these conditions, it also carries the standing re-certification a +generation deeper — this time against **two** recorded artifact sets at once +(details in D19). + +## Design extraction (verbatim, source-cited) + +**The bound M4 closes — M3-BRIEF, Honest limits (frozen with M3's artifacts):** +"One bound worth stating plainly: **the matrix measures collateral among 12 +concepts, not across the vocabulary.** M1 established breadth over 60 concepts +with one control each; M3 establishes near-zero collateral over 132 ordered +pairs of 12. Nothing here shows that deleting France spares the other 48 M1 +concepts — that is a different (and cheap) experiment the matrix design +deliberately did not run." + +**The scope as HANDOFF recorded it (option 4):** "A 12-prime × 60-probe strip +would close that gap for roughly the cost of one M3 subject run." (The honest +cell count is larger than that quote — ~4.3× an M1 subject run, still +minutes-scale and $0; wall-clock section below.) + +**The prediction under test — M3-BRIEF, finding 1:** "Collateral concentrates +on *probes*, not on *primes* … at 1.5B all 11 off-diagonal misses land on +`silver` (6), `Canada` (3), `piano` (1) and `violin` (1); at 3B all 9 land on +`silver` (5) and four others once each." + +**The machinery carried unchanged:** D9(b) oracle (`oracle.py`, byte-shared), +D16's grade-recorded-cells-first re-certification pattern, D17's wide-oracle +degeneracy guard and verdict precedence, D18's two run-time bars, M2's frozen +late thirds (L17–21 / L19–24 / L26–32), per-item `direction_key` as recorded. +Every strip cell ablates the subject's identical late-third layer set at λ = 1, +k = 1, so cells differ only in which direction is removed and which item is +asked — the M3 property, preserved (PR #7 F2's caveat clause carries). + +**Instrument facts the design stands on (computed 2026-07-28, from the +recorded artifacts and the frozen `oracle.py` — no new model runs):** + +- **Gated sets are already known, exactly.** Gating is the clean arm under + D9(b), deterministic and recorded. Full-roster prefix-gated n: **69 / 105 / + 116** (0.5B / 1.5B / 3B) of 180 items. Non-subset gated items — the gate + arm's n: **41 / 71 / 84**. Non-subset concepts with ≥ 1 gated item: **23 / + 41 / 43** of 48 (zero-gated: 25 / 7 / 5). Those concepts drop out because + **the model answers something else, or answers correctly in a form D9(b)'s + opening-word rule refuses** — never because of token geometry: under + D9(b) the gate is a prefix on the 3-token span, so it is insensitive to + token count up to 3, and four of the seven 1.5B zero-gated concepts are + single-token bare. The recorded clean spans say it plainly — `saturn-1/2/3 + → 'Jupiter'`, `neptune-1 → 'Jupiter'`, `bronze-2 → 'Aluminum'`, + `monday-1/2/3 → 'Friday'`, `thursday-1/3 → 'Friday'`, `moth-1 → 'Wasp'`. + Two further mechanisms sit beside that one and are named rather than folded + in: `ant-1/2 → 'Ants'` is a **morphology** miss (a plural closes no + boundary), and D9(b)'s opening-word rule refuses **correct answers behind a + modifier** — `duck-3 → 'Peking duck'` and `beetle-3 → 'Insect Beetle'` at 3B, + `eagle-1 → 'American eagle'`, `guitar-2 → 'Electric guitar'` and + `bear-3 → 'Polar bear'` at 0.5B. All three mechanisms point the same way: + the gate arm is **the concepts this subject already names, in the form the + oracle accepts** — a *competence selection*, and confidently-named concepts + are plausibly the robust ones, which biases the measured sparing floor + **upward**. That is the same enrichment mechanism as F12's + selection-enrichment finding, now on the probe side; it is why the claim is + sparing across the *measurable* vocabulary, said exactly that way. +- **Five probe clues mention a prime's spelling — a confound M3's design never + had.** M1's leak guard (D5) bars a clue from leaking its *own* concept or + control; M3 verified no cross-mentions *within the 12*. Widening probes to + 180 items surfaces five (prime, item) pairs, scanned with **D5's own rule** — + no word of the clue may *start with* the string, case-insensitive, plus that + string's `forbidden_forms` entries — not the narrower whole-word match: + **October→september-2, silver→flute-1, China→jade-1, October→opal-2**, and + **Egypt→beetle-2** ("Ancient **Egyptians** carved amulets of the scarab, one + kind of this insect"), which a whole-word scan misses because the clue + inflects the prime. In those cells a miss cannot distinguish "collateral + damage to naming machinery" from "the clue's own text lost a word it + references." The first four items gate at both gate-bearing subjects (only + jade-1 gates at 0.5B), so those cells are live; `beetle-2` is **ungated on + all three subjects** (clean span 'Scarab' / 'scarab'), so it contributes no + gate-bearing cell today and is listed so a future re-gate cannot silently + miss it. D21 decides their treatment. +- **D9(b)'s owned span-truncation residual re-enters a gate-bearing arm for + the first time since M1.** `oracle.py`'s frozen wording owns a residual: for + the three concepts whose bare spelling *fills* the 3-token span — beetle, + butterfly, trumpet — the closing boundary is unobservable, so a longer word + sharing those tokens would score as a hit ("Beetlejuice" truncates to + exactly "Beetle"). The wording's closing scope sentence, "**None of the + three concepts is in M2's subset**," is precisely what made the residual + harmless for M2 and M3: those concepts were never scored. M4 widens the + probe side to all 60, so their items re-enter a **gate-bearing** pool. + The residual condition is exact and narrower than "a 3-token word": the + recorded span, after stripping leading whitespace, **equals the concept's + spelling with nothing following it**, compared case-insensitively as + `oracle.says_concept_prefix` compares, so no boundary character is observed. + ("Fills the 3-token span" would not distinguish anything — every recorded + `greedy_3` is exactly 3 tokens by construction; and a case-*exact* + comparison would select nothing, since the spans are capitalised and the + roster spellings are not.) On that reading the gate-arm residual cells are + **0 / 2 / 2** at 0.5B / 1.5B / 3B — 1.5B `beetle-1` ('Beetle') and + `butterfly-1` ('Butterfly') of 71; 3B `trumpet-3` ('Trumpet') and + `butterfly-1` ('Butterfly') of 84; none gated at 0.5B. These are the + gate-arm members of the same residual set `oracle.py`'s frozen docstring + names — the docstring counts **six recorded M1 cells** ("beetle-1, + butterfly-1 ×2 arms, trumpet-3"), enumerating across arms (`butterfly-1` + in both `clean` and `control_late`), where the four above are clean-arm + cells on the two gate-bearing subjects. The other trumpet cells are *not* + residual: `trumpet-1` at 1.5B and `trumpet-1` / `trumpet-2` at 3B record + `'Trumpet<|im_end|>'`, and `<|im_end|>` closes a word under + `oracle._BOUNDARY` — capitalisation is why, since generated `Trumpet` is two + tokens (`['Trump','et']`) where bare `trumpet` is three, leaving room for the + terminator. D22's span bar cannot catch the residual: it passes at exactly + ≤ 3 tokens, which *is* the residual condition. The bias runs **toward** the + gate — an unobservably-terminated span scores a hit, inflating survival — + and at 1.5B the pass/fail margin is one item (44/71 passes, 43/71 fails), so + even two such items exceed the margin. `oracle.py` stays untouched (editing + it would force re-runs of three milestones); the residual is carried by + disclosure plus the pre-registered residual-conservative recomputation in + D20. +- **The D18 bars pass for the whole roster, checked in advance.** All 60 + concepts tokenize to ≤ `oracle.SPAN_TOKENS` = 3 in both bare and + leading-space form on all three Qwen2.5 tokenizers (verified 2026-07-28), + and all 60 spellings are pure ASCII. The bars still run at run time per + D18's rationale — the check above is design due-diligence, not a substitute. +- **What the recorded evidence already fixes — and what it predicts.** + Thirteen concepts' frozen M1 controls are subset members, so their recorded + `control_late` cells are strip cells. The non-subset portion is a + single-direction, *same-category* proxy for the new pool — the least + favorable cell type, since M3 found what collateral exists sits within + category at small scale: it reads **6/8** at 1.5B and **13/15** at 3B. + Being strip cells, these are **in-sample**, not a forecast: their outcomes + are already determined, so the misses among them **cap the gate arm before a + single new forward pass** — pre-registered ceilings of **35/41, 69/71 and + 82/84** at 0.5B / 1.5B / 3B (0.5B: all 6 proxy cells miss; 1.5B: `july-3` + and `venus-3` miss; 3B: `guitar-2` and `neptune-1` miss). At 1.5B the bar + needs 44 of 71, so the ceiling leaves real room; at 3B likewise. Where the + proxy is genuinely predictive is the *rest* of the pool, which is mostly + cross-category — so a high floor is the prediction at the gate-bearing + subjects. +- **The 0.5B prediction, read in M4's own statistic.** The strongest and most + directly comparable evidence is not the 6-cell proxy: it is M3's recorded + subset, recomputed under **M4's gate statistic**. M3's per-item + survives-all-11 off-diagonal reads **19/28 = 0.679 at 0.5B, Wilson lower + bound 0.4934** — *below 0.5*, i.e. **M3's own 0.5B subset would already fail + M4's bar**. The same field passes at both gate-bearing subjects (1.5B 29/34, + lower 0.6987; 3B 27/32, lower 0.6825). So there is no "tension" between M3's + clean 0.5B floor and a possible M4 failure to resolve — M3's 0.5B floor was + clean only under the *cluster-mean per-cell* statistic, and switching to the + conjunction flips it. That is also the sharpest illustration of why the 0.5 + constant carries no meaning across statistics (D20). The 0.5B strip may well + fail its floor — off-gate, under the standing any-direction-damage frame, + and consistent in advance with M3's own numbers, M1's full-battery 0.5B + control cell (33/69), the 0/6 proxy, and F12's selection-enrichment finding + (the subset-12's 0.5B robustness came partly from S1's selection rule). + +## Dispositions for the two inherited obligations (explicit, per HANDOFF) + +**(1) The `oracle._BOUNDARY` boundary-class decision (D18's named trigger).** +Not owed: the trigger is a stage freezing a **non-ASCII** list, and M4 adds no +vocabulary — every probe and every prime comes from M1's frozen 60, all +pure-ASCII spellings. The premise stays pinned, not assumed: D18's ASCII bar +runs in the M4 runner's pre-trial validation, unchanged. `oracle.py` is +untouched (it would become byte-shared by a fourth consumer — deviations +table). The trigger continues to name the banked scope stretch (idea #13), not +this stage. + +**(2) The conjunction-degeneracy rule (PR #9 F1's carry-forward).** The rule: +any stage whose gate is a conjunction must put every surviving-side comparison +arm on the dispositive degeneracy list and enumerate them in its frozen +wording. Every gate option in D20 is deliberately **single-clause**, so the +dispositive list has exactly one surviving arm — the pooled non-subset +off-target cell — and D20's wording names it explicitly. Stated affirmatively +so the obligation reads as discharged by design, not forgotten: if a later +amendment ever makes M4's gate a conjunction, every surviving arm goes on the +list before the wording freezes. + +## Decisions to freeze (Kyle picks; recommendations flagged) + +*Frozen (Kyle, 2026-07-29): **D19 (a)** the full 12 × 180 strip plus a full +clean re-run (2,340 cells/subject) with the embedded 255-cell M1 and 468-cell +M3 re-certifications graded first; **D20 (a)** the single-clause +survives-all-12 level gate as written, with the verdict-label choice resolved +per review F8 — the label stays `VOCAB-SPARING` and the **realized survival +proportion rides inside the verdict string**, the M3 `ON A DAMAGED FLOOR` +pattern applied to a level bar; **D21 (a)** all five cross-mention pairs kept +in the gate-bearing pool with named per-cell reporting; **D22 (a)** both +run-time instrument bars over all 60 scored words plus the 12 direction words. +Full `DECISIONS.md` entries (D19–D22) land with the M4 code PR, per the +M0/M1/M2/M3 pattern. Amended pre-freeze at PR #10's adversarial review, across +three rounds: F1–F4 (should-fix) fixed and verified; F11–F12 (should-fix, +defects of the F2 fix) fixed and verified; F16 (should-fix, a defect of the +F11 fix — a case-exact selector that would have silently matched zero cells) +fixed with F15/F17/F18 at `edb5387`; and all eight deferred nice-to-haves +(F5–F7, F9, F10, F13–F15) pulled in at the freeze on Kyle's call ("pull in all +8"), because F5/F7/F9/F13's text byte-freezes into code once these decisions +are signed. Round 4 was authorized by Kyle beyond the three-dispatch cap to +verify the `edb5387` fixes and this freeze commit.* + +*Amended post-freeze, pre-run, at round 4 (F20 + F21) — the M3 precedent for a +review-driven amendment before any cell is run, **ratified by Kyle 2026-07-29** +("I ratify the F20 amendment"): the verdict string as first frozen carried +only the **as-scored** proportion, so the two pre-registered reads that can flip +which number is honest (residual-conservative fail-in-place; concept-level +collapse) stayed in prose — reproducing the exact failure F8 was resolved to +prevent. D20's wording now carries the pre-declared **AS-SCORED ONLY** +qualifier, attached conditionally by the runner whenever a conservative read's +Wilson lower bound falls below 0.5 while the as-scored read's does not, and +names a failing label that the single pass-label template had left unstated. +The gate, its 0.5 bar, its arm and its precondition are unchanged — this scopes +the claim and, per D17's carried rule, can never create or rescue one.* + +***Amendment 2, post-freeze, pre-run, at rounds 5–6 (F22, F24–F27) —*** +***ratified by Kyle 2026-07-29*** *("I agree with what you recommend"), on the +recommendation that named all four changes below. A second material change to +the same frozen +`GATE_WORDING` block, recorded separately rather than folded into Amendment 1, +because Kyle's ratification quote covers Amendment 1 only. What changed:* **(i) +the failing label** *is now the lineage's pre-committed null* `not shown` +*rather than the assertive* `NOT VOCAB-SPARING` *Amendment 1 introduced — +failing a Wilson* lower *bound cannot establish the contrary (at 1.5B, k = 40 +has a point estimate of 0.563* above *the bar with a straddling interval), and +all three predecessor runners emit* `not shown`*;* **(ii) 0.5B is scoped inside +the wording** *— the gate verdict is the AND over the two gate-bearing subjects +and 0.5B's readout is never a gate claim, where Amendment 1 had pre-committed +the string "per subject" while the same block declares 0.5B never gate-bearing;* +**(iii) the qualifier's attachment is restricted** *to a claim-level verdict, +never to* `NOT A RESULT` / `DEGENERATE` / `UNDERPOWERED` *— verbatim, Amendment +1 attached it to all of them, contradicting the D17 rule it cites (the fix +copies* `m3_matrix.py`*'s own docstring rule); and* **(iv) both string templates +are stated explicitly** *with a fixed read order, where Amendment 1 gave the +qualifier an example but no template. The gate, its 0.5 bar, its arm and its +re-certification precondition remain unchanged — verified byte-for-byte against +`90b994c`.* + +### D19 — Primes × probes: the strip frame (decide first) + +- **(a) 12 subset primes × all 180 M1 items, plus a full clean re-run + (recommended).** Per subject: `clean` (180 items) + 12 × 180 = 2,160 ablated + cells = **2,340 cells**. *Why:* + - **The new claim gets its cells** — every gated non-subset item under every + characterized direction (492 / 852 / 1,008 cells). + - **The re-certification surface is maximal, and two generations deep.** The + strip contains **255** cells per subject recorded in M1's artifacts + (`clean` 180, the 12 subset concepts' `primed_late` 36, and the 39 + `control_late` cells whose control direction is a subset member) **and + 468** cells recorded in M3's artifacts (the 36 subset `clean` cells + all + 432 matrix cells — everything M3 ran except its 18 out-of-subset + control-extras, whose directions are not strip primes). Both comparisons + are on raw recorded strings (`greedy`, `greedy_3`; `concept_mass` as + texture), graded first, INVALID on mismatch on the certified stack — the + D16 pattern applied against two artifact sets at once. + - **Ungated items ride along as texture** (the standing convention — the + gate reads only gated cells, but every cell is recorded). +- **(b) 12 primes × the 144 non-subset items only.** Saves ~20% of the run and + destroys the M3-overlap re-certification (no subset cells, no diagonal) plus + the in-strip recorded proxies. The one check that has caught nothing yet + *because it runs every time* — broken for one saved coffee break. Not + recommended. +- **(c) The full 60 × 180 matrix.** Answers a different, bigger question ("is + *every* direction safe to delete?") at ~4.7× the cost, with 48 primes whose + collateral behaviour no milestone has characterized and no pre-registered + expectation exists for. That is a future stage's question (it shares rails + with banked idea #13), not this close-out's. Not recommended. + +### D20 — The pre-committed wording package (gate, degeneracy, precedence) + +**Gate options (the substance Kyle picks; wording frozen as code in +`m4_strip.GATE_WORDING` before any run, written verbatim into every results +JSON):** + +- **(a) A survives-everything level gate on the new pool (recommended).** + + > **VOCAB-SPARING** iff, per subject: among the gated **non-subset** items + > (concepts outside the 12-concept matrix roster), the proportion that + > **survives all 12** subset-direction deletions — D9(b) naming success in + > every one of the item's 12 off-target cells — has its Wilson 95% lower + > bound at or above **0.5**. This 0.5 is a bar on the **12-fold + > conjunction**, not on per-cell survival, and is **not** M3's per-cell + > floor: under independence it corresponds to a per-cell survival of + > 0.5^(1/12) ≈ **0.944**, and its stringency depends on the deletion count + > (12) as much as on the sparing rate. The bar is read **only when the 468 + > M3-recorded and 255 M1-recorded cells in the strip reproduce their recorded + > outcomes bit-for-bit**; any mismatch is INVALID and there is no verdict. + > The M4 verdict is the AND over 1.5B and 3B; 0.5B runs and is reported under + > its standing any-direction-damage frame, never gate-bearing. Gate-arm + > n < MIN_N = 20 ⇒ pre-declared UNDERPOWERED and no claim (realized n = + > 41 / 71 / 84 from the recorded gated sets). **Verdict string, + > pre-committed:** the label alone over-reads — clearing a floor bar is + > compatible with a large minority of measurable items damaged — so the + > verdict, whichever way it goes, carries its **realized survival proportion + > in the same string**: `VOCAB-SPARING` when the bar is cleared and the + > lineage's pre-committed null **`not shown`** when it is not — never an + > assertive negative, because failing a Wilson *lower* bound does not + > establish the contrary (`m1_battery.py`, `m2_depth.py`, `m3_matrix.py` all + > emit `not shown`). The gate verdict is the AND over the two gate-bearing + > subjects; 0.5B's readout is reported in the same shape under its standing + > any-direction-damage frame and is **not** a gate verdict, so a low 0.5B + > reading is never a `not shown` gate claim. + > **Conservative-read qualifier, pre-declared:** two pre-registered reads can + > fall below the bar when the as-scored read clears it — the + > residual-conservative fail-in-place read and the concept-level collapse — + > and this brief pre-commits that where they diverge, *their* numbers are the + > honest ones. So a **claim-level** verdict additionally carries the qualifier + > **AS-SCORED ONLY**, naming each such read and its number, whenever any + > pre-registered conservative read's Wilson 95% lower bound is below **0.5** + > while the as-scored read's is not. The qualifier **scopes a claim and can + > never create or rescue one** (D17's rule, carried), so it attaches to a + > bar-level verdict only and never to `NOT A RESULT`, `DEGENERATE` or + > `UNDERPOWERED` — precedence has already withheld the claim there, leaving + > it nothing to scope. + > **The two templates, stated once and implemented verbatim** — base: + > `