From c42dbf3b23797c2c9c4e4ad47749e06b2e35371e Mon Sep 17 00:00:00 2001 From: ksdisch Date: Tue, 28 Jul 2026 19:29:51 -0500 Subject: [PATCH 01/13] =?UTF-8?q?docs(m4):=20start-of-stage=20brief=20?= =?UTF-8?q?=E2=80=94=20vocabulary=20collateral=20strip,=20D19=E2=80=93D22?= =?UTF-8?q?=20for=20freeze?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit M4 added by Kyle's post-M3 pick (2026-07-28): close M3's stated bound (collateral among 12, not across the vocabulary) before write-up + /seed-hunt. S1/S2 stretches declined and banked as idea #13 in the j-lens-proj-ideas backlog. Brief contents: 12 primes x 180 probes frame (D19), single-clause VOCAB-SPARING level gate on the never-measured non-subset pool (D20), treatment of the four cross-mention confound cells found while drafting (D21), D18 bars widened to all 60 scored words (D22). Realized ns computed from recorded artifacts with the frozen oracle; both inherited obligations dispositioned explicitly. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01ANtkrPFv6i9CgCasojqcoZ --- HANDOFF.md | 39 +++--- docs/M4-BRIEF.md | 352 +++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 372 insertions(+), 19 deletions(-) create mode 100644 docs/M4-BRIEF.md diff --git a/HANDOFF.md b/HANDOFF.md index 50c183d..24870f8 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -1,6 +1,6 @@ # HANDOFF.md — mute-map -_Last updated: 2026-07-28_ +_Last updated: 2026-07-28 (post-M3 decision session)_ ## What was just done @@ -78,24 +78,25 @@ and each needs its own brief. ## Immediate next move -**Nothing is pre-committed.** The honest options, none of them owed: - -1. **Close the project** — a write-up of the four-milestone arc, then - `/seed-hunt` for the next paper. The characterization KICKOFF bought is - delivered: breadth (M1), localization + dose (M2), specificity (M3). -2. **S1 stretch — scale.** 7B lens fit on a rented GPU (≤ $15, decision K3's - no-refit rule applies only to the core chain) plus a matrix-lite. The - sharpening-with-scale story is the one M1–M3 keep gesturing at and never - measured above 3B. -3. **S2 stretch — scope.** Token mute button or concept mute button? - (translations, synonyms, morphological variants). **Its brief owes the - `oracle._BOUNDARY` boundary-class decision before it freezes any non-ASCII - list** — that is the named future trigger D18 recorded, and it is the only - inherited obligation on the board. -4. **The cheap follow-up M3 explicitly did not run:** M3 measures collateral - among 12 concepts, not across the vocabulary. Nothing here shows that deleting - France spares the other 48 M1 concepts. A 12-prime × 60-probe strip would - close that gap for roughly the cost of one M3 subject run. +**Decided (Kyle, 2026-07-28, session "mute-map post M3 decisions"): run the +vocabulary collateral strip as close-out stage M4, then write-up + +`/seed-hunt`.** The S1/S2 stretches were declined for this repo and banked as +idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md` — they +compete in the seed-hunt on equal terms, no incumbent's privilege. + +**Where M4 stands: `docs/M4-BRIEF.md` is written and in review; decisions +D19–D22 await Kyle's freeze.** The stage: 12 subset primes × all 180 M1 items +(2,340 cells/subject), gate proposed as a single-clause VOCAB-SPARING level +bar on the never-measured non-subset pool (per-item survives-all-12, Wilson +lower bound ≥ 0.5 at 1.5B AND 3B, the 0.5 carried from M3's pre-registered +floor). Realized ns are known from the recorded gated sets (gate arm 41 / 71 / +84). Design facts found while drafting: four probe clues mention a prime's +spelling (October→september-2, silver→flute-1, China→jade-1, October→opal-2 — +D21 decides their treatment); all 60 roster words pass the D18 span bar on all +three tokenizers (checked in advance); the recorded same-category proxy for +the new pool reads 6/8 (1.5B) and 13/15 (3B) but **0/6 at 0.5B**, so the +off-gate 0.5B floor may genuinely fail — reportable under the standing frame. +No runner code exists yet; code only after freeze, cut from `m3_matrix.py`. Standing constraints unchanged: certified environment = `mps` + torch 2.13.0 + transformers 5.13.1 (off it: NOT A RESULT); `m0_anchor.py` stays certified and diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md new file mode 100644 index 0000000..3d2d12b --- /dev/null +++ b/docs/M4-BRIEF.md @@ -0,0 +1,352 @@ +# M4 start-of-stage brief — the vocabulary collateral strip + +*Start-of-stage brief per the per-stage rhythm: plain-terms explanation first, +design extraction second, decisions third, code only after Kyle freezes. +Decisions here are D19–D22, continuing `docs/DECISIONS.md` (D15–D18 were M3's).* + +*Provenance, stated up front: M4 is **not** in KICKOFF's frozen chain. The v1 +chain (M0–M3) closed on 2026-07-28; this stage is the close-out follow-up +HANDOFF listed as option 4, picked by Kyle the same day (session "mute-map post +M3 decisions") over closing immediately and over the S1/S2 stretches — which +were banked as idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md`. +KICKOFF's scope decisions are not relitigated; this is an addition, owned in the +deviations table. After M4, the plan of record is write-up + `/seed-hunt`.* + +## What M4 is, in plain terms + +M3's killer figure shows a dark diagonal on a nearly-white grid: deleting any +of the 12 subset concepts' directions mutes that concept and almost nothing +else — *among those 12*. M3's own Honest limits section states the bound +plainly: nothing measured so far shows that deleting France spares the other +48 M1 concepts. The natural over-reading of the figure ("the deletion spares +the vocabulary") is exactly one experiment wider than what was run. + +M4 runs that experiment. Keep the 12 characterized directions as the **primes** +(the thing deleted); widen the **probes** (the thing asked about) from the 12 +subset concepts to **all 60 M1 concepts** — the full frozen 180-item battery. +Every cell is the M3 recipe unchanged: ablate prime A's direction at the +subject's late third (λ = 1, k = 1), ask item-of-B its naming question, score +with the frozen D9(b) oracle. The strip is 12 × 180 = 2,160 ablated cells per +subject; the genuinely new content is the **non-subset pool** — the 48 concepts' +gated items under each of the 12 deletions, cells no milestone has ever +measured (492 / 852 / 1,008 of them at 0.5B / 1.5B / 3B). + +The strip is also an out-of-sample test of M3's most interesting descriptive +finding. Finding 1 said collateral concentrates on a few fragile **probes** +(silver's column) rather than on damaging **primes**. The 48 new probe columns +are a fresh sample: either near-total sparing holds and the fragile-probe story +stays a curiosity of the 12, or more silver-like columns exist among the 48 — +both outcomes are findings, and a *failed* gate (a low survival floor across +the wider vocabulary) would be the pre-committed reportable headline against +the specificity story's reach. + +Because the strip contains, cell-for-cell, almost everything M1 and M3 already +recorded on these conditions, it also carries the standing re-certification a +generation deeper — this time against **two** recorded artifact sets at once +(details in D19). + +## Design extraction (verbatim, source-cited) + +**The bound M4 closes — M3-BRIEF, Honest limits (frozen with M3's artifacts):** +"One bound worth stating plainly: **the matrix measures collateral among 12 +concepts, not across the vocabulary.** M1 established breadth over 60 concepts +with one control each; M3 establishes near-zero collateral over 132 ordered +pairs of 12. Nothing here shows that deleting France spares the other 48 M1 +concepts — that is a different (and cheap) experiment the matrix design +deliberately did not run." + +**The scope as HANDOFF recorded it (option 4):** "A 12-prime × 60-probe strip +would close that gap for roughly the cost of one M3 subject run." (The honest +cell count is larger than that quote — ~4.3× an M1 subject run, still +minutes-scale and $0; wall-clock section below.) + +**The prediction under test — M3-BRIEF, finding 1:** "Collateral concentrates +on *probes*, not on *primes* … at 1.5B all 11 off-diagonal misses land on +`silver` (6), `Canada` (3), `piano` (1) and `violin` (1); at 3B all 9 land on +`silver` (5) and four others once each." + +**The machinery carried unchanged:** D9(b) oracle (`oracle.py`, byte-shared), +D16's grade-recorded-cells-first re-certification pattern, D17's wide-oracle +degeneracy guard and verdict precedence, D18's two run-time bars, M2's frozen +late thirds (L17–21 / L19–24 / L26–32), per-item `direction_key` as recorded. +Every strip cell ablates the subject's identical late-third layer set at λ = 1, +k = 1, so cells differ only in which direction is removed and which item is +asked — the M3 property, preserved (PR #7 F2's caveat clause carries). + +**Instrument facts the design stands on (computed 2026-07-28, from the +recorded artifacts and the frozen `oracle.py` — no new model runs):** + +- **Gated sets are already known, exactly.** Gating is the clean arm under + D9(b), deterministic and recorded. Full-roster prefix-gated n: **69 / 105 / + 116** (0.5B / 1.5B / 3B) of 180 items. Non-subset gated items — the gate + arm's n: **41 / 71 / 84**. Non-subset concepts with ≥ 1 gated item: **23 / + 41 / 43** of 48 (zero-gated: 25 / 7 / 5 — mostly the bare-multi-token words + D9 documented; their absence from the gate arm is tokenizer geometry, owned + since M1). +- **Four probe clues mention a prime's spelling — a confound M3's design never + had.** M1's leak guard (D5) bars a clue from leaking its *own* concept or + control; M3 verified no cross-mentions *within the 12*. Widening probes to + 180 items surfaces exactly four (prime, item) pairs where the item's clue + contains the deleted word at a word boundary: **October→september-2, + silver→flute-1, China→jade-1, October→opal-2**. In those cells a miss cannot + distinguish "collateral damage to naming machinery" from "the clue's own + text lost a word it references." All four items gate at both gate-bearing + subjects (only jade-1 gates at 0.5B), so the cells are live. D21 decides + their treatment. +- **The D18 bars pass for the whole roster, checked in advance.** All 60 + concepts tokenize to ≤ `oracle.SPAN_TOKENS` = 3 in both bare and + leading-space form on all three Qwen2.5 tokenizers (verified 2026-07-28), + and all 60 spellings are pure ASCII. The bars still run at run time per + D18's rationale — the check above is design due-diligence, not a substitute. +- **What the recorded evidence predicts.** Thirteen concepts' frozen M1 + controls are subset members, so their recorded `control_late` cells are + strip cells. The non-subset portion is a single-direction, *same-category* + proxy for the new pool — the least favorable cell type, since M3 found + what collateral exists sits within category at small scale: it reads **6/8** + at 1.5B and **13/15** at 3B. The pool itself is mostly cross-category, so a + high floor is the prediction at the gate-bearing subjects. At 0.5B the same + proxy reads **0/6** — consistent with M1's full-battery 0.5B control cell + (33/69) and with F12's selection-enrichment finding (the subset-12's 0.5B + robustness came partly from S1's selection rule), and *not* with M3's clean + 0.5B subset floor. The strip may well fail its floor at 0.5B — off-gate, + under the standing any-direction-damage frame, and reportable as the + measured resolution of that tension. + +## Dispositions for the two inherited obligations (explicit, per HANDOFF) + +**(1) The `oracle._BOUNDARY` boundary-class decision (D18's named trigger).** +Not owed: the trigger is a stage freezing a **non-ASCII** list, and M4 adds no +vocabulary — every probe and every prime comes from M1's frozen 60, all +pure-ASCII spellings. The premise stays pinned, not assumed: D18's ASCII bar +runs in the M4 runner's pre-trial validation, unchanged. `oracle.py` is +untouched (it would become byte-shared by a fourth consumer — deviations +table). The trigger continues to name the banked scope stretch (idea #13), not +this stage. + +**(2) The conjunction-degeneracy rule (PR #9 F1's carry-forward).** The rule: +any stage whose gate is a conjunction must put every surviving-side comparison +arm on the dispositive degeneracy list and enumerate them in its frozen +wording. Every gate option in D20 is deliberately **single-clause**, so the +dispositive list has exactly one surviving arm — the pooled non-subset +off-target cell — and D20's wording names it explicitly. Stated affirmatively +so the obligation reads as discharged by design, not forgotten: if a later +amendment ever makes M4's gate a conjunction, every surviving arm goes on the +list before the wording freezes. + +## Decisions to freeze (Kyle picks; recommendations flagged) + +### D19 — Primes × probes: the strip frame (decide first) + +- **(a) 12 subset primes × all 180 M1 items, plus a full clean re-run + (recommended).** Per subject: `clean` (180 items) + 12 × 180 = 2,160 ablated + cells = **2,340 cells**. *Why:* + - **The new claim gets its cells** — every gated non-subset item under every + characterized direction (492 / 852 / 1,008 cells). + - **The re-certification surface is maximal, and two generations deep.** The + strip contains **255** cells per subject recorded in M1's artifacts + (`clean` 180, the 12 subset concepts' `primed_late` 36, and the 39 + `control_late` cells whose control direction is a subset member) **and + 468** cells recorded in M3's artifacts (the 36 subset `clean` cells + all + 432 matrix cells — everything M3 ran except its 18 out-of-subset + control-extras, whose directions are not strip primes). Both comparisons + are on raw recorded strings (`greedy`, `greedy_3`; `concept_mass` as + texture), graded first, INVALID on mismatch on the certified stack — the + D16 pattern applied against two artifact sets at once. + - **Ungated items ride along as texture** (the standing convention — the + gate reads only gated cells, but every cell is recorded). +- **(b) 12 primes × the 144 non-subset items only.** Saves ~20% of the run and + destroys the M3-overlap re-certification (no subset cells, no diagonal) plus + the in-strip recorded proxies. The one check that has caught nothing yet + *because it runs every time* — broken for one saved coffee break. Not + recommended. +- **(c) The full 60 × 180 matrix.** Answers a different, bigger question ("is + *every* direction safe to delete?") at ~4.7× the cost, with 48 primes whose + collateral behaviour no milestone has characterized and no pre-registered + expectation exists for. That is a future stage's question (it shares rails + with banked idea #13), not this close-out's. Not recommended. + +### D20 — The pre-committed wording package (gate, degeneracy, precedence) + +**Gate options (the substance Kyle picks; wording frozen as code in +`m4_strip.GATE_WORDING` before any run, written verbatim into every results +JSON):** + +- **(a) A survives-everything level gate on the new pool (recommended).** + + > **VOCAB-SPARING** iff, per subject: among the gated **non-subset** items + > (concepts outside the 12-concept matrix roster), the proportion that + > **survives all 12** subset-direction deletions — D9(b) naming success in + > every one of the item's 12 off-target cells — has its Wilson 95% lower + > bound at or above **0.5**. The M4 verdict is the AND over 1.5B and 3B; + > 0.5B runs and is reported under its standing any-direction-damage frame, + > never gate-bearing. Gate-arm n (41 / 71 / 84) < MIN_N = 20 ⇒ pre-declared + > UNDERPOWERED and no claim. + + *Why this shape.* The strip's question is a **level** question — "is the + floor high?" — not an ordering question; M3 already settled the ordering. + The per-item survives-all-12 outcome is a true binary, so the Wilson + interval is exact for it — this deliberately does **not** promote M3's + cluster-mean floor readout to gate-bearing, because D17 froze that + approximation as "acceptable only because the qualifier is never + dispositive," and M4 keeps that rationale intact (the cluster-mean floor is + reported beside, reference line 0.5, never dispositive — the M3-comparable + view). The **0.5 constant is carried from M3's pre-registered floor**, + deliberately not fitted to any recorded cell. It is a floor bar, not an + effect-size claim — the descriptive numbers carry the actual size. + *Trade-off, owned:* survives-all-12 is the strictest sparing statistic; a + single fragile cell fails an item, and the correlation structure across the + 12 deletions (unmeasured until this run) decides how harsh that is. That + bias direction runs **against** the claim, the one direction this project + accepts. + +- **(b) An M3-clause-(1)-style ordering gate extended to the strip** (pooled + off-target minus subset diagonal, Newcombe CI excludes 0). Maximally + comparable to M3 — and it passes almost by inheritance, since the diagonal + is 0-to-3-hits at every subject and the off-target pool would have to + collapse to near-zero to close a Newcombe gap that large. A gate the + recorded evidence has effectively already decided does not gate the new + claim. Not recommended (reported beside as descriptive continuity either + way). + +- **(c) The conjunction of (a) AND (b).** Strongest-sounding wording; adds + nothing (b) doesn't already concede, and re-opens the conjunction + obligation for no inferential gain. Not recommended. + +**Degeneracy disposition (D14/D17's wide-oracle guard, re-scoped to M4's +arms).** The dispositive guard is unchanged in mechanism: pool the first +tokens of an arm's **non-produced** cells only, share against the arm's full +cell count, threshold COLLAPSE_SHARE = 0.5. Scope, enumerated: collapse in the +pooled **non-subset off-target** arm — the single surviving arm the gate +reads — ⇒ **DEGENERATE**, no VOCAB-SPARING claim; collapse in the subset +**diagonal** ⇒ **TAG only** (the expected mute signature, carried); `clean` +stays off the dispositive list (the D14 F3 correction, carried); collapse +inside any single prime's row, any probe concept's column, or any per-pair +cell is **texture**, attached to the readout it compromises. + +**The effective-n sanity check, pre-registered and never dispositive.** Items +cluster three-per-concept on the probe side (they share the concept whose +fragility is being measured), so beside the item-level gate the same statistic +is recomputed collapsed to one binary per **concept** — "every gated item of +this concept survives all 12 deletions" — over the non-subset concepts with +≥ 1 gated item (n = **23 / 41 / 43**, all ≥ MIN_N). If the item-level gate is +clean and the concept-level collapse is not, the concept-level numbers are the +honest ones to quote. + +**Descriptive package, never gate-bearing, all pre-registered here:** the +cluster-mean per-cell floor on the new pool (M3's F15 readout, reference line +0.5, the M3-comparable view); the ordering contrast from option (b); per-prime +**row profiles** (does any of the 12 damage the wider vocabulary?) and +per-probe **column profiles** over all 60 concepts (finding 1's out-of-sample +test — are there silver-like columns among the 48?); the within- vs +cross-category split of the new pool; the four confound cells (D21); mean +concept mass per cell under D13's standing scope; the 0.5B floor under its +standing frame (the recorded proxy predicts it may fail — that outcome is a +finding, not a failure). + +**Verdict precedence, frozen** in `strip_verdict()` (`m4_verdict.py`): NOT A +RESULT > DEGENERATE > UNDERPOWERED > the level bar. Wrong-arm inputs exit +INVALID before the checkpoint loads (the D18-companion shape, carried from +`m3_matrix.py`); `--dry-run` validates and stops; `--limit` is smoke, never a +result; M4 refuses M1 or M3 artifacts that were themselves not results. + +### D21 — The four cross-mention cells + +- **(a) Keep them in the gate-bearing pool; report them as a named confound + row (recommended).** The four cells stay in every pooled arm and in their + items' survives-all-12 conjunctions, and the results section reports each + cell's outcome individually under its named confound. *Why:* a confounded + miss can only **lower** the floor — the bias runs against the gate, the one + direction this project ships owned. Excluding them would delete only cells + that could hurt the claim, which is the anti-conservative move the lineage + never makes. Four cells of 852 (1.5B) cannot carry a verdict either way; + what they can do is mislead a *reader* of the column profiles, and the named + row prevents that. +- **(b) Pre-registered exclusion from gate-bearing pools, reported as + texture.** Cleaner causal story per cell, but it is evidence-removal in the + gate's favour — rejected on the standing bias rule. Not recommended. +- **(c) Drop the four items entirely.** Loses their clean cells and their 11 + unconfounded prime cells for no reason. Not recommended. + +### D22 — Run-time instrument bars (D18 carried, widened to the full roster) + +- **(a) Both D18 bars in `m4_strip.py`'s pre-trial validation, now over every + scored word — all 60 — plus the 12 direction words, with unit tests + (recommended).** Span bar: `max(len(tok(w)), len(tok(" " + w))) ≤ + SPAN_TOKENS` on the subject's own tokenizer, else INVALID. ASCII bar: every + spelling pure ASCII, else INVALID. Verified in advance for all 60 words on + all three tokenizers (2026-07-28, instrument facts above) — the run-time bar + still runs, because D18's point was that the premise must hold *at the + moment of measurement*. `oracle.py` untouched. +- **(b) Bars over the 12 primes only (M3's literal scope).** The soundness + premise attaches to every **scored** word, and M4 scores 60 — a bar that + checks 12 of them pins a fifth of the premise. Not recommended. +- **(c) Widen `_BOUNDARY` now.** Still zero live cases; still the unforced + version of the mistake D9 exists to prevent. Not recommended (carried + rejection). + +## Deviations table additions (owned) + +| Deviation | From | Owned reason | +|---|---|---| +| M4 exists at all (post-KICKOFF stage) | KICKOFF's frozen M0–M3 chain | Kyle-picked close-out (2026-07-28) that closes M3's stated bound before write-up; KICKOFF's scope decisions unrelitigated; S1/S2 declined and banked (idea #13) | +| Level-bar gate (Wilson lower bound vs a constant) | the lineage's Newcombe ordering gates | The ordering is M3's settled result; the strip's question is a level question; the 0.5 constant is carried from M3's pre-registered floor, not fitted | +| Four cross-mention (prime, item) cells kept in gate-bearing pools | M3's verified no-cross-mention property | M1's leak guard only bars own-concept/control leaks; the confound biases against the gate; named per-cell reporting (D21a) | +| `oracle.py` byte-shared by a fourth consumer (`m4_strip.py`) | cut-from-predecessor rule | Same D9 rationale as the first three consumers: the rule's purpose is byte-identity; pinned by the existing shared-oracle test pattern | +| Probe-side reach is still the D9(b)-visible roster | "the vocabulary" | 25 / 7 / 5 concepts gate zero items (tokenizer geometry, owned since M1); the claim is sparing across the *measurable* vocabulary, said exactly that way | + +Standing owned rows carry unchanged: naming-only gate (K2), lens provenance +(K3), S2-stratum space-keyed directions (D11), mass-channel scope (D13), +identical-layer-set arms (D17's carried clause). + +## Expected power (honest math — realized, not projected) + +Gating is the deterministic clean arm, already recorded, so every n is known +now (a run that disagrees is an INVALID cross-check, not a power surprise): + +| Cell | 0.5B | 1.5B | 3B | Clears MIN_N = 20? | +|---|---|---|---|---| +| Gated items, full roster | 69 | 105 | 116 | yes | +| **Gate arm: gated non-subset items (per-item survives-all-12)** | **41** | **71** | **84** | yes, all | +| New off-target cells (non-subset gated × 12) | 492 | 852 | 1,008 | yes (texture pools) | +| Concept-level collapse (non-subset concepts ≥ 1 gated item) | 23 | 41 | 43 | yes, all | +| Subset diagonal (recorded; M3's cells re-run) | 28 | 34 | 32 | yes | +| Confound cells in the pool (D21a) | 1 | 4 | 4 | named texture | + +Honesty rows, carried and extended: probe-side clustering (3 items share a +concept) is handled by the pre-registered concept-level collapse, never by the +gate silently; prime-side clustering (every item faces the same 12 directions) +is the unmeasured correlation structure the survives-all-12 statistic is +conservative under; the pooled 852-cell view repeats each item 12 times and is +therefore texture, never the gate (M3's F11 lesson applied in advance). The +worked bar: at 1.5B the gate passes iff ≥ **44 of 71** items survive all 12 +deletions (wilson(44, 71) lower bound ≈ 0.503, computed with the project's own +frozen ruler; 43/71 reads 0.489 and fails); the recorded proxies and M3's +per-item collapse texture (29/34 survived all 11 at 1.5B) predict clearance +with room — and at 0.5B the 0/6 proxy predicts the off-gate floor may fail, +which would be the first measured divergence between the subset's 0.5B +robustness and the wider roster's. + +## Wall-clock plan + +Per subject under D19(a): 180 clean + 2,160 ablated = **2,340 cells × 3 +forwards ≈ 7,020 forwards — ~4.3× an M1 subject run** (1,620), which ran in +minutes. All three subjects comfortably within an afternoon on MPS, $0, +backgrounded with untracked logs. The 255 M1-recorded and 468 M3-recorded +cells are graded first (D19a); wrong-arm inputs exit INVALID before the +checkpoint loads; `--dry-run` validates and stops; `--limit` is smoke, never a +result; `m4_strip.GATE_WORDING` frozen as code before any real run. Runner cut +from `m3_matrix.py` (never from certified predecessors), verdict in +`m4_verdict.py`, tests in `test_m4.py`. + +## What M4 does NOT decide + +- **How the project closes** — the write-up + `/seed-hunt` flow is the plan of + record after M4 and is not a milestone decision. +- **The banked stretches** (idea #13: scope + scale) — declined for this repo; + they compete on equal terms in the seed-hunt. +- **M1's, M2's and M3's published verdicts and numbers** — they stand as + pre-committed; the strip's subset cells re-certify them, never re-litigate + them. +- **No oracle change.** `oracle.py` stays byte-identical in all four + consumers; D22 pins its premises without altering the rule. Never an LLM + judge, never free-text parsing (standing guardrail). From 98eecbd99bc4799b34058e822be5d4d9acf62779 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Tue, 28 Jul 2026 20:23:22 -0500 Subject: [PATCH 02/13] =?UTF-8?q?docs(m4):=20review=20round=201=20fixes=20?= =?UTF-8?q?=E2=80=94=20F1=E2=80=93F4?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit F1: own the 0.5 gate constant as new, uncalibrated and deliberately lenient rather than "carried from M3" — the statistic changed (12-fold conjunction vs per-cell; 0.5 on the conjunction ≈ 0.944 per-cell under independence) and the status changed (D17 qualifier that could never rescue a claim → sole dispositive gate). The per-cell equivalence and the deletion-count dependency go into GATE_WORDING itself; new deviations row. F2: own D9(b)'s span-truncation residual, which re-enters a gate-bearing arm for the first time since M1 now that probes widen to all 60 — 0 / 3 / 4 span-filling gate-arm items (1.5B trumpet-1, beetle-1, butterfly-1; 3B trumpet-1/2/3, butterfly-1), biasing toward the gate where the 1.5B margin is one item. Adds an instrument fact, a power-table row, a deviations row and a pre-registered residual-conservative recomputation (never dispositive), and corrects the claim that D22 pins the oracle's premises — it pins the span premise, not the scope premise. oracle.py untouched. F3: restate zero-gating (25 / 7 / 5 concepts) as competence-selection, not tokenizer geometry — under D9(b) the gate is token-count-insensitive up to 3 and the recorded clean spans are wrong answers, not truncations. States the upward bias on the floor and connects it to F12's selection-enrichment. F4: correct "cells no milestone has ever measured" to 486 / 844 / 993 (of 492 / 852 / 1,008); the remaining 6 / 8 / 15 are M1-recorded — the very cells used as the proxy — so they are in-sample and cap the gate arm at 35/41, 69/71, 82/84 before any forward pass. Power-table row split, ceilings stated as pre-registered fact, out-of-sample framing tempered. HANDOFF.md mirrors the F1 and F4 corrections. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- HANDOFF.md | 11 +-- docs/M4-BRIEF.md | 171 ++++++++++++++++++++++++++++++++++++----------- 2 files changed, 140 insertions(+), 42 deletions(-) diff --git a/HANDOFF.md b/HANDOFF.md index 24870f8..b92463b 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -87,10 +87,13 @@ compete in the seed-hunt on equal terms, no incumbent's privilege. **Where M4 stands: `docs/M4-BRIEF.md` is written and in review; decisions D19–D22 await Kyle's freeze.** The stage: 12 subset primes × all 180 M1 items (2,340 cells/subject), gate proposed as a single-clause VOCAB-SPARING level -bar on the never-measured non-subset pool (per-item survives-all-12, Wilson -lower bound ≥ 0.5 at 1.5B AND 3B, the 0.5 carried from M3's pre-registered -floor). Realized ns are known from the recorded gated sets (gate arm 41 / 71 / -84). Design facts found while drafting: four probe clues mention a prime's +bar on the non-subset pool (per-item survives-all-12, Wilson lower bound ≥ 0.5 +at 1.5B AND 3B — a *new*, deliberately lenient, uncalibrated constant on the +12-fold conjunction, ≈ 0.944 per-cell under independence, explicitly not +carried from M3's per-cell floor). Realized ns are known from the recorded +gated sets (gate arm 41 / 71 / 84); 486 / 844 / 993 of the pool's 492 / 852 / +1,008 cells are genuinely new, the rest are M1-recorded and cap the gate arm at +35/41, 69/71, 82/84 before any forward pass. Design facts found while drafting: four probe clues mention a prime's spelling (October→september-2, silver→flute-1, China→jade-1, October→opal-2 — D21 decides their treatment); all 60 roster words pass the D18 span bar on all three tokenizers (checked in advance); the recorded same-category proxy for diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md index 3d2d12b..3ba9522 100644 --- a/docs/M4-BRIEF.md +++ b/docs/M4-BRIEF.md @@ -28,14 +28,20 @@ Every cell is the M3 recipe unchanged: ablate prime A's direction at the subject's late third (λ = 1, k = 1), ask item-of-B its naming question, score with the frozen D9(b) oracle. The strip is 12 × 180 = 2,160 ablated cells per subject; the genuinely new content is the **non-subset pool** — the 48 concepts' -gated items under each of the 12 deletions, cells no milestone has ever -measured (492 / 852 / 1,008 of them at 0.5B / 1.5B / 3B). - -The strip is also an out-of-sample test of M3's most interesting descriptive -finding. Finding 1 said collateral concentrates on a few fragile **probes** -(silver's column) rather than on damaging **primes**. The 48 new probe columns -are a fresh sample: either near-total sparing holds and the fragile-probe story -stays a curiosity of the 12, or more silver-like columns exist among the 48 — +gated items under each of the 12 deletions: **492 / 852 / 1,008** cells at +0.5B / 1.5B / 3B, of which **486 / 844 / 993** have never been measured by any +milestone. The remainder (6 / 8 / 15) are already recorded in M1 — they are the +non-subset items whose frozen M1 control direction happens to be one of the 12 +primes — so those outcomes are fixed before the run and are stated below as +pre-registered ceilings, not predictions. + +The strip is also a mostly out-of-sample test of M3's most interesting +descriptive finding. Finding 1 said collateral concentrates on a few fragile +**probes** (silver's column) rather than on damaging **primes**. The 48 new +probe columns are a near-fresh sample — 6 / 8 / 15 of their gate-bearing cells +are M1-recorded, the other 486 / 844 / 993 are new: either near-total sparing +holds and the fragile-probe story stays a curiosity of the 12, or more +silver-like columns exist among the 48 — both outcomes are findings, and a *failed* gate (a low survival floor across the wider vocabulary) would be the pre-committed reportable headline against the specificity story's reach. @@ -80,9 +86,20 @@ recorded artifacts and the frozen `oracle.py` — no new model runs):** D9(b), deterministic and recorded. Full-roster prefix-gated n: **69 / 105 / 116** (0.5B / 1.5B / 3B) of 180 items. Non-subset gated items — the gate arm's n: **41 / 71 / 84**. Non-subset concepts with ≥ 1 gated item: **23 / - 41 / 43** of 48 (zero-gated: 25 / 7 / 5 — mostly the bare-multi-token words - D9 documented; their absence from the gate arm is tokenizer geometry, owned - since M1). + 41 / 43** of 48 (zero-gated: 25 / 7 / 5). Those concepts drop out because + **the model names something else**, not because of token geometry: under + D9(b) the gate is a prefix on the 3-token span, so it is insensitive to + token count up to 3, and four of the seven 1.5B zero-gated concepts are + single-token bare. The recorded clean spans say it plainly — `saturn-1/2/3 + → 'Jupiter'`, `neptune-1 → 'Jupiter'`, `bronze-2 → 'Aluminum'`, + `monday-1/2/3 → 'Friday'`, `thursday-1/3 → 'Friday'`, `moth-1 → 'Wasp'`; + `ant-1/2 → 'Ants'` is a morphology miss (a plural closes no boundary). + So the gate arm is **the concepts this subject already names correctly** — a + *competence selection*, and confidently-named concepts are plausibly the + robust ones, which biases the measured sparing floor **upward**. That is the + same enrichment mechanism as F12's selection-enrichment finding, now on the + probe side; it is why the claim is sparing across the *measurable* + vocabulary, said exactly that way. - **Four probe clues mention a prime's spelling — a confound M3's design never had.** M1's leak guard (D5) bars a clue from leaking its *own* concept or control; M3 verified no cross-mentions *within the 12*. Widening probes to @@ -93,17 +110,44 @@ recorded artifacts and the frozen `oracle.py` — no new model runs):** text lost a word it references." All four items gate at both gate-bearing subjects (only jade-1 gates at 0.5B), so the cells are live. D21 decides their treatment. +- **D9(b)'s owned span-truncation residual re-enters a gate-bearing arm for + the first time since M1.** `oracle.py`'s frozen wording owns a residual: for + the three concepts whose bare spelling *fills* the 3-token span — beetle, + butterfly, trumpet — the closing boundary is unobservable, so a longer word + sharing those tokens would score as a hit ("Beetlejuice" truncates to + exactly "Beetle"). The wording's closing scope sentence, "**None of the + three concepts is in M2's subset**," is precisely what made the residual + harmless for M2 and M3: those concepts were never scored. M4 widens the + probe side to all 60, so their items re-enter a **gate-bearing** pool: + **0 / 3 / 4** span-filling gate-arm items at 0.5B / 1.5B / 3B — 1.5B + `trumpet-1`, `beetle-1`, `butterfly-1` (of 71); 3B `trumpet-1`, `trumpet-2`, + `trumpet-3`, `butterfly-1` (of 84); none gated at 0.5B. D22's span bar + cannot catch this: it passes at exactly ≤ 3 tokens, which *is* the residual + condition. The bias runs **toward** the gate — an unobservably-terminated + span scores a hit, inflating survival — and at 1.5B the pass/fail margin is + one item (44/71 passes, 43/71 fails), so three such items exceed the margin. + `oracle.py` stays untouched (editing it would force re-runs of three + milestones); the residual is carried by disclosure plus the pre-registered + residual-conservative recomputation in D20. - **The D18 bars pass for the whole roster, checked in advance.** All 60 concepts tokenize to ≤ `oracle.SPAN_TOKENS` = 3 in both bare and leading-space form on all three Qwen2.5 tokenizers (verified 2026-07-28), and all 60 spellings are pure ASCII. The bars still run at run time per D18's rationale — the check above is design due-diligence, not a substitute. -- **What the recorded evidence predicts.** Thirteen concepts' frozen M1 - controls are subset members, so their recorded `control_late` cells are - strip cells. The non-subset portion is a single-direction, *same-category* - proxy for the new pool — the least favorable cell type, since M3 found - what collateral exists sits within category at small scale: it reads **6/8** - at 1.5B and **13/15** at 3B. The pool itself is mostly cross-category, so a +- **What the recorded evidence already fixes — and what it predicts.** + Thirteen concepts' frozen M1 controls are subset members, so their recorded + `control_late` cells are strip cells. The non-subset portion is a + single-direction, *same-category* proxy for the new pool — the least + favorable cell type, since M3 found what collateral exists sits within + category at small scale: it reads **6/8** at 1.5B and **13/15** at 3B. + Being strip cells, these are **in-sample**, not a forecast: their outcomes + are already determined, so the misses among them **cap the gate arm before a + single new forward pass** — pre-registered ceilings of **35/41, 69/71 and + 82/84** at 0.5B / 1.5B / 3B (0.5B: all 6 proxy cells miss; 1.5B: `july-3` + and `venus-3` miss; 3B: `guitar-2` and `neptune-1` miss). At 1.5B the bar + needs 44 of 71, so the ceiling leaves real room; at 3B likewise. Where the + proxy is genuinely predictive is the *rest* of the pool. The pool itself is + mostly cross-category, so a high floor is the prediction at the gate-bearing subjects. At 0.5B the same proxy reads **0/6** — consistent with M1's full-battery 0.5B control cell (33/69) and with F12's selection-enrichment finding (the subset-12's 0.5B @@ -177,10 +221,14 @@ JSON):** > (concepts outside the 12-concept matrix roster), the proportion that > **survives all 12** subset-direction deletions — D9(b) naming success in > every one of the item's 12 off-target cells — has its Wilson 95% lower - > bound at or above **0.5**. The M4 verdict is the AND over 1.5B and 3B; - > 0.5B runs and is reported under its standing any-direction-damage frame, - > never gate-bearing. Gate-arm n (41 / 71 / 84) < MIN_N = 20 ⇒ pre-declared - > UNDERPOWERED and no claim. + > bound at or above **0.5**. This 0.5 is a bar on the **12-fold + > conjunction**, not on per-cell survival, and is **not** M3's per-cell + > floor: under independence it corresponds to a per-cell survival of + > 0.5^(1/12) ≈ **0.944**, and its stringency depends on the deletion count + > (12) as much as on the sparing rate. The M4 verdict is the AND over 1.5B + > and 3B; 0.5B runs and is reported under its standing + > any-direction-damage frame, never gate-bearing. Gate-arm n (41 / 71 / 84) + > < MIN_N = 20 ⇒ pre-declared UNDERPOWERED and no claim. *Why this shape.* The strip's question is a **level** question — "is the floor high?" — not an ordering question; M3 already settled the ordering. @@ -188,11 +236,33 @@ JSON):** interval is exact for it — this deliberately does **not** promote M3's cluster-mean floor readout to gate-bearing, because D17 froze that approximation as "acceptable only because the qualifier is never - dispositive," and M4 keeps that rationale intact (the cluster-mean floor is - reported beside, reference line 0.5, never dispositive — the M3-comparable - view). The **0.5 constant is carried from M3's pre-registered floor**, - deliberately not fitted to any recorded cell. It is a floor bar, not an - effect-size claim — the descriptive numbers carry the actual size. + dispositive," and M4 keeps *that* rationale intact for the approximation + (the cluster-mean floor is reported beside, reference line 0.5, never + dispositive — the M3-comparable view). What does **not** carry is the + constant itself; see immediately below. + + *The 0.5 constant is new, and owned as new — not carried from M3.* M3's 0.5 + was a floor on cluster-collapsed **per-cell** survival; M4's is a bar on a + **12-fold conjunction**. The same digits mean opposite things across those + two statistics: under independence a per-cell 0.5 corresponds to a + conjunction of 0.5^12 ≈ 0.0002, and a conjunction of 0.5 corresponds to a + per-cell 0.5^(1/12) ≈ 0.944. Two consequences, stated rather than inherited. + **(i) Status changed.** M3's constant was itself *uncalibrated* — M3-BRIEF's + F14 record: "the pooled off-diagonal has never been measured at any scale … + which is the honest reason the constant is uncalibrated" — and D17 tolerated + that only because the qualifier it scoped "can never create or rescue" a + claim. M4 makes a constant of the same value the **single dispositive gate**, + so D17's tolerance does not transfer either. **(ii) The deletion count is + half the bar.** Stringency is set as much by *how many* deletions each item + must survive (12) as by the sparing rate: at a per-cell rate of 0.971 (M3's + recorded 1.5B off-diagonal) the conjunction reads ≈ 0.70 and clears the bar; + at 0.94 it reads ≈ 0.48 and fails. M4's 0.5 is therefore a **new, + deliberately lenient, uncalibrated constant** — pre-registered here before + any new cell is run, and fitted to no recorded cell — and its per-cell + equivalence is written into `GATE_WORDING` itself so no write-up can quote it + as M3's floor. It is a floor bar, not an effect-size claim; the descriptive + numbers carry the actual size. + *Trade-off, owned:* survives-all-12 is the strictest sparing statistic; a single fragile cell fails an item, and the correlation structure across the 12 deletions (unmeasured until this run) decides how harsh that is. That @@ -232,12 +302,25 @@ this concept survives all 12 deletions" — over the non-subset concepts with clean and the concept-level collapse is not, the concept-level numbers are the honest ones to quote. +**The residual-conservative read, pre-registered and never dispositive.** +Beside the gate as scored, the same gate statistic is recomputed with every +scored cell whose recorded greedy span *fills* the 3-token window re-scored as +a **miss** — the maximally conservative reading of the boundary D9(b) cannot +observe. In the recorded clean arm those are the beetle / butterfly / trumpet +items listed in the instrument facts (0 / 3 / 4 gate-arm items); in the +ablated cells the set is whatever the run records, computed from the same +recorded spans. Both numbers are reported. Following the same honesty pattern +as the concept-level collapse: **if the gate passes and the +residual-conservative read does not, the conservative numbers are the honest +ones to quote.** The read can only lower the floor, never rescue it, and +`oracle.py` is not touched. + **Descriptive package, never gate-bearing, all pre-registered here:** the cluster-mean per-cell floor on the new pool (M3's F15 readout, reference line 0.5, the M3-comparable view); the ordering contrast from option (b); per-prime **row profiles** (does any of the 12 damage the wider vocabulary?) and -per-probe **column profiles** over all 60 concepts (finding 1's out-of-sample -test — are there silver-like columns among the 48?); the within- vs +per-probe **column profiles** over all 60 concepts (finding 1's mostly +out-of-sample test — are there silver-like columns among the 48?); the within- vs cross-category split of the new pool; the four confound cells (D21); mean concept mass per cell under D13's standing scope; the 0.5B floor under its standing frame (the recorded proxy predicts it may fail — that outcome is a @@ -289,10 +372,12 @@ result; M4 refuses M1 or M3 artifacts that were themselves not results. | Deviation | From | Owned reason | |---|---|---| | M4 exists at all (post-KICKOFF stage) | KICKOFF's frozen M0–M3 chain | Kyle-picked close-out (2026-07-28) that closes M3's stated bound before write-up; KICKOFF's scope decisions unrelitigated; S1/S2 declined and banked (idea #13) | -| Level-bar gate (Wilson lower bound vs a constant) | the lineage's Newcombe ordering gates | The ordering is M3's settled result; the strip's question is a level question; the 0.5 constant is carried from M3's pre-registered floor, not fitted | +| Level-bar gate (Wilson lower bound vs a constant) | the lineage's Newcombe ordering gates | The ordering is M3's settled result; the strip's question is a level question | +| A **new, uncalibrated, sole-dispositive** 0.5 constant | D17's 0.5, which was per-cell and never dispositive | The statistic changed (12-fold conjunction, ≈ 0.944 per-cell under independence) and the status changed (qualifier → gate), so no provenance transfers; owned as deliberately lenient, pre-registered before any new cell, fitted to none, with the per-cell equivalence frozen into `GATE_WORDING` | | Four cross-mention (prime, item) cells kept in gate-bearing pools | M3's verified no-cross-mention property | M1's leak guard only bars own-concept/control leaks; the confound biases against the gate; named per-cell reporting (D21a) | | `oracle.py` byte-shared by a fourth consumer (`m4_strip.py`) | cut-from-predecessor rule | Same D9 rationale as the first three consumers: the rule's purpose is byte-identity; pinned by the existing shared-oracle test pattern | -| Probe-side reach is still the D9(b)-visible roster | "the vocabulary" | 25 / 7 / 5 concepts gate zero items (tokenizer geometry, owned since M1); the claim is sparing across the *measurable* vocabulary, said exactly that way | +| D9(b)'s owned span-truncation residual sits in a gate-bearing arm | `oracle.py`'s scope sentence, "None of the three concepts is in M2's subset" | M4 scores all 60 probes, so beetle / butterfly / trumpet are gated again for the first time since M1 (0 / 3 / 4 gate-arm items); the bias runs *toward* the gate, so it is disclosed per subject and carried by the pre-registered residual-conservative recomputation (D20), never by editing `oracle.py` | +| Probe-side reach is still the D9(b)-visible roster | "the vocabulary" | 25 / 7 / 5 concepts gate zero items — a **competence selection** (the model names something else; the gate is token-count-insensitive up to 3), not tokenizer geometry; that selection plausibly enriches for robust concepts and biases the floor **upward** (F12's enrichment mechanism, probe side); the claim is sparing across the *measurable* vocabulary, said exactly that way | Standing owned rows carry unchanged: naming-only gate (K2), lens provenance (K3), S2-stratum space-keyed directions (D11), mass-channel scope (D13), @@ -307,10 +392,13 @@ now (a run that disagrees is an INVALID cross-check, not a power surprise): |---|---|---|---|---| | Gated items, full roster | 69 | 105 | 116 | yes | | **Gate arm: gated non-subset items (per-item survives-all-12)** | **41** | **71** | **84** | yes, all | -| New off-target cells (non-subset gated × 12) | 492 | 852 | 1,008 | yes (texture pools) | +| Off-target cells in the new pool (non-subset gated × 12) | 492 | 852 | 1,008 | yes (texture pools) | +| — of them never measured by any milestone | 486 | 844 | 993 | texture | +| — of them already recorded in M1 (outcome fixed pre-run) | 6 | 8 | 15 | ceilings 35/41, 69/71, 82/84 | | Concept-level collapse (non-subset concepts ≥ 1 gated item) | 23 | 41 | 43 | yes, all | | Subset diagonal (recorded; M3's cells re-run) | 28 | 34 | 32 | yes | | Confound cells in the pool (D21a) | 1 | 4 | 4 | named texture | +| Span-filling residual items in the gate arm (D9(b), clean arm) | 0 | 3 | 4 | named texture | Honesty rows, carried and extended: probe-side clustering (3 items share a concept) is handled by the pre-registered concept-level collapse, never by the @@ -320,11 +408,14 @@ conservative under; the pooled 852-cell view repeats each item 12 times and is therefore texture, never the gate (M3's F11 lesson applied in advance). The worked bar: at 1.5B the gate passes iff ≥ **44 of 71** items survive all 12 deletions (wilson(44, 71) lower bound ≈ 0.503, computed with the project's own -frozen ruler; 43/71 reads 0.489 and fails); the recorded proxies and M3's -per-item collapse texture (29/34 survived all 11 at 1.5B) predict clearance -with room — and at 0.5B the 0/6 proxy predicts the off-gate floor may fail, -which would be the first measured divergence between the subset's 0.5B -robustness and the wider roster's. +frozen ruler; 43/71 reads 0.489 and fails). Between 44 and that arm's +pre-registered ceiling of **69 of 71** — 8 of the 852 cells are M1-recorded and +two of them (`july-3`, `venus-3`) already miss — there is real room, and M3's +per-item collapse texture (29/34 survived all 11 at 1.5B) points into it; the +ceiling is a fact about the arm, not evidence for the gate. At 0.5B the same +recorded cells put the ceiling at **35 of 41** (all 6 proxy cells miss), so the +off-gate floor may well fail there, which would be the first measured +divergence between the subset's 0.5B robustness and the wider roster's. ## Wall-clock plan @@ -348,5 +439,9 @@ from `m3_matrix.py` (never from certified predecessors), verdict in pre-committed; the strip's subset cells re-certify them, never re-litigate them. - **No oracle change.** `oracle.py` stays byte-identical in all four - consumers; D22 pins its premises without altering the rule. Never an LLM - judge, never free-text parsing (standing guardrail). + consumers; D22 pins its ≤ 3-token span premise without altering the rule. + What D22 cannot pin is the wording's *scope* premise — "None of the three + concepts is in M2's subset" — which M4's wider probe side retires. That + residual is owned in the deviations table and carried by the + residual-conservative read (D20), not by a rule change. Never an LLM judge, + never free-text parsing (standing guardrail). From 35a8743fa2e7e05d6f4e234b2cd375cb5579c5c6 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Tue, 28 Jul 2026 20:35:58 -0500 Subject: [PATCH 03/13] =?UTF-8?q?docs(m4):=20review=20round=202=20fixes=20?= =?UTF-8?q?=E2=80=94=20F11,=20F12?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both are defects of the F2 fix, not of the finding it addressed. F11: the residual-conservative selector was degenerate and its counts were wrong. "The span fills the 3-token window" matches every recorded cell — greedy_3 IS the 3-token span. The real condition is narrower: the span, after stripping leading whitespace, equals the concept's spelling with nothing following it, so no boundary is observed. On that reading the gate-arm residual is 0 / 2 / 2, not 0 / 3 / 4 — 1.5B beetle-1, butterfly-1; 3B trumpet-3, butterfly-1 — which is exactly the set oracle.py's frozen docstring already names ("beetle-1, butterfly-1 x2 arms, trumpet-3"). trumpet-1 (1.5B) and trumpet-1/-2 (3B) record 'Trumpet<|im_end|>' and <|im_end|> closes a word under oracle._BOUNDARY, so they carry no residual; generated 'Trumpet' is two tokens where bare 'trumpet' is three, which leaves room for the terminator. Corrected in all three places (instrument fact, power table, deviations row). F12: the conservative read's denominator was left to the runner, and the two readings straddle the bar (same numerator: wilson(43,71) lower 0.489 fails, wilson(43,69) lower 0.505 passes). Pre-registers fail-in-place — the arm stays at its as-scored n (41 / 71 / 84, preserving the power table's cross-check) and residual-affected items score as failures — and names and rejects the un-gating alternative as the less conservative of the two. F13-F15 (nice-to-have) remain FOLLOW-UP pending Kyle's pull-in call. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- docs/M4-BRIEF.md | 69 +++++++++++++++++++++++++++++++++--------------- 1 file changed, 47 insertions(+), 22 deletions(-) diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md index 3ba9522..9f00313 100644 --- a/docs/M4-BRIEF.md +++ b/docs/M4-BRIEF.md @@ -118,17 +118,29 @@ recorded artifacts and the frozen `oracle.py` — no new model runs):** exactly "Beetle"). The wording's closing scope sentence, "**None of the three concepts is in M2's subset**," is precisely what made the residual harmless for M2 and M3: those concepts were never scored. M4 widens the - probe side to all 60, so their items re-enter a **gate-bearing** pool: - **0 / 3 / 4** span-filling gate-arm items at 0.5B / 1.5B / 3B — 1.5B - `trumpet-1`, `beetle-1`, `butterfly-1` (of 71); 3B `trumpet-1`, `trumpet-2`, - `trumpet-3`, `butterfly-1` (of 84); none gated at 0.5B. D22's span bar - cannot catch this: it passes at exactly ≤ 3 tokens, which *is* the residual - condition. The bias runs **toward** the gate — an unobservably-terminated - span scores a hit, inflating survival — and at 1.5B the pass/fail margin is - one item (44/71 passes, 43/71 fails), so three such items exceed the margin. - `oracle.py` stays untouched (editing it would force re-runs of three - milestones); the residual is carried by disclosure plus the pre-registered - residual-conservative recomputation in D20. + probe side to all 60, so their items re-enter a **gate-bearing** pool. + The residual condition is exact and narrower than "a 3-token word": the + recorded span, after stripping leading whitespace, **equals the concept's + spelling with nothing following it**, so no boundary character is observed. + ("Fills the 3-token span" would not distinguish anything — every recorded + `greedy_3` is exactly 3 tokens by construction.) On that reading the + gate-arm residual cells are **0 / 2 / 2** at 0.5B / 1.5B / 3B — 1.5B + `beetle-1` ('Beetle') and `butterfly-1` ('Butterfly') of 71; 3B `trumpet-3` + ('Trumpet') and `butterfly-1` ('Butterfly') of 84; none gated at 0.5B. This + is exactly the set `oracle.py`'s frozen docstring already names ("beetle-1, + butterfly-1 ×2 arms, trumpet-3"). The other trumpet cells are *not* + residual: `trumpet-1` at 1.5B and `trumpet-1` / `trumpet-2` at 3B record + `'Trumpet<|im_end|>'`, and `<|im_end|>` closes a word under + `oracle._BOUNDARY` — capitalisation is why, since generated `Trumpet` is two + tokens (`['Trump','et']`) where bare `trumpet` is three, leaving room for the + terminator. D22's span bar cannot catch the residual: it passes at exactly + ≤ 3 tokens, which *is* the residual condition. The bias runs **toward** the + gate — an unobservably-terminated span scores a hit, inflating survival — + and at 1.5B the pass/fail margin is one item (44/71 passes, 43/71 fails), so + even two such items exceed the margin. `oracle.py` stays untouched (editing + it would force re-runs of three milestones); the residual is carried by + disclosure plus the pre-registered residual-conservative recomputation in + D20. - **The D18 bars pass for the whole roster, checked in advance.** All 60 concepts tokenize to ≤ `oracle.SPAN_TOKENS` = 3 in both bare and leading-space form on all three Qwen2.5 tokenizers (verified 2026-07-28), @@ -304,15 +316,28 @@ honest ones to quote. **The residual-conservative read, pre-registered and never dispositive.** Beside the gate as scored, the same gate statistic is recomputed with every -scored cell whose recorded greedy span *fills* the 3-token window re-scored as -a **miss** — the maximally conservative reading of the boundary D9(b) cannot -observe. In the recorded clean arm those are the beetle / butterfly / trumpet -items listed in the instrument facts (0 / 3 / 4 gate-arm items); in the -ablated cells the set is whatever the run records, computed from the same -recorded spans. Both numbers are reported. Following the same honesty pattern -as the concept-level collapse: **if the gate passes and the -residual-conservative read does not, the conservative numbers are the honest -ones to quote.** The read can only lower the floor, never rescue it, and +**residual cell** re-scored as a **miss** — the maximally conservative reading +of the boundary D9(b) cannot observe. A residual cell is one whose recorded +span, after stripping leading whitespace, **equals the scored concept's +spelling exactly, with nothing following it** (no boundary character observed). +That is the selector, stated once here and implemented verbatim: it is *not* +"the span fills the 3-token window", which every recorded cell does. In the +recorded clean arm the residual cells are the 0 / 2 / 2 gate-arm items named in +the instrument facts; in the ablated cells the set is whatever the run records, +computed from the same recorded spans by the same selector. + +**Denominator, pre-registered: fail in place.** The gate arm stays at its +as-scored n — **41 / 71 / 84**, the knowable-now cross-check the power table +freezes — and a residual-affected item scores as a **failure** within that arm. +The alternative reading (re-score the clean cell too, so the item un-gates and +the arm shrinks) is named and **rejected**: it is the less conservative of the +two at the bar — with the same numerator, `wilson(43, 71)` reads 0.489 and +fails while `wilson(43, 69)` reads 0.505 and passes — and it would break the +power table's pre-registered n. Fail-in-place keeps the read strictly one-way: +it can only lower the floor, never rescue it. Both the as-scored and the +conservative numbers are reported, and following the same honesty pattern as +the concept-level collapse: **if the gate passes and the residual-conservative +read does not, the conservative numbers are the honest ones to quote.** `oracle.py` is not touched. **Descriptive package, never gate-bearing, all pre-registered here:** the @@ -376,7 +401,7 @@ result; M4 refuses M1 or M3 artifacts that were themselves not results. | A **new, uncalibrated, sole-dispositive** 0.5 constant | D17's 0.5, which was per-cell and never dispositive | The statistic changed (12-fold conjunction, ≈ 0.944 per-cell under independence) and the status changed (qualifier → gate), so no provenance transfers; owned as deliberately lenient, pre-registered before any new cell, fitted to none, with the per-cell equivalence frozen into `GATE_WORDING` | | Four cross-mention (prime, item) cells kept in gate-bearing pools | M3's verified no-cross-mention property | M1's leak guard only bars own-concept/control leaks; the confound biases against the gate; named per-cell reporting (D21a) | | `oracle.py` byte-shared by a fourth consumer (`m4_strip.py`) | cut-from-predecessor rule | Same D9 rationale as the first three consumers: the rule's purpose is byte-identity; pinned by the existing shared-oracle test pattern | -| D9(b)'s owned span-truncation residual sits in a gate-bearing arm | `oracle.py`'s scope sentence, "None of the three concepts is in M2's subset" | M4 scores all 60 probes, so beetle / butterfly / trumpet are gated again for the first time since M1 (0 / 3 / 4 gate-arm items); the bias runs *toward* the gate, so it is disclosed per subject and carried by the pre-registered residual-conservative recomputation (D20), never by editing `oracle.py` | +| D9(b)'s owned span-truncation residual sits in a gate-bearing arm | `oracle.py`'s scope sentence, "None of the three concepts is in M2's subset" | M4 scores all 60 probes, so beetle / butterfly / trumpet are gated again for the first time since M1 (0 / 2 / 2 gate-arm cells whose span equals the spelling with no boundary observed — `oracle.py`'s own named set); the bias runs *toward* the gate, so it is disclosed per subject and carried by the pre-registered residual-conservative recomputation, fail-in-place (D20), never by editing `oracle.py` | | Probe-side reach is still the D9(b)-visible roster | "the vocabulary" | 25 / 7 / 5 concepts gate zero items — a **competence selection** (the model names something else; the gate is token-count-insensitive up to 3), not tokenizer geometry; that selection plausibly enriches for robust concepts and biases the floor **upward** (F12's enrichment mechanism, probe side); the claim is sparing across the *measurable* vocabulary, said exactly that way | Standing owned rows carry unchanged: naming-only gate (K2), lens provenance @@ -398,7 +423,7 @@ now (a run that disagrees is an INVALID cross-check, not a power surprise): | Concept-level collapse (non-subset concepts ≥ 1 gated item) | 23 | 41 | 43 | yes, all | | Subset diagonal (recorded; M3's cells re-run) | 28 | 34 | 32 | yes | | Confound cells in the pool (D21a) | 1 | 4 | 4 | named texture | -| Span-filling residual items in the gate arm (D9(b), clean arm) | 0 | 3 | 4 | named texture | +| D9(b) residual items in the gate arm (clean arm; span = spelling, no boundary) | 0 | 2 | 2 | named texture | Honesty rows, carried and extended: probe-side clustering (3 items share a concept) is handled by the pre-registered concept-level collapse, never by the From 9f5749bba4e69371f0694b06bd6b5b43c13282c1 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Tue, 28 Jul 2026 21:50:27 -0500 Subject: [PATCH 04/13] =?UTF-8?q?docs(m4):=20pull=20in=20all=20eight=20rev?= =?UTF-8?q?iew=20follow-ups=20=E2=80=94=20F5-F7,=20F9,=20F10,=20F13-F15?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Kyle's call at the freeze gate: pull all eight in before D19-D22 freeze, since four of them get byte-frozen into code once he signs. F5: re-scan the cross-mention confound with D5's OWN rule (prefix + forbidden_forms, not whole-word) — five pairs, not four. Adds Egypt->beetle-2 ("Ancient Egyptians carved amulets of the scarab..."), which a whole-word scan misses because the clue inflects the prime. beetle-2 is ungated on all three subjects ('Scarab'/'scarab'), so it carries no gate-bearing cell today; it is listed so a future re-gate cannot silently acquire one. D21 retitled, list frozen in its text, deviations row and power-table row updated. F6: replace the proxy-only 0.5B prediction with the in-statistic number. M3's recorded 0.5B subset reads 19/28 = 0.679 under M4's own gate statistic, Wilson lower 0.4934 — it ALREADY FAILS M4's bar. So there is no "tension" between M3's clean 0.5B floor and an M4 failure: M3's floor was clean only under the cluster-mean per-cell statistic, and switching to the conjunction flips it. Sharpest illustration of D20's point that the constant carries no meaning across statistics. F7: move the realized ns outside the underpowered conditional — the proposed wording literally read "41 / 71 / 84 < 20". Now abstract, with realized n in a trailing parenthetical, matching m1_battery.GATE_WORDING's shape. F9: add the re-certification precondition to GATE_WORDING. The single-clause level bar, read alone, is satisfiable by a no-op intervention (~100% survival prints VOCAB-SPARING). The guarantee existed in D19's design and the runner's exit code but not in the sentence a write-up quotes; now the sentence carries its own precondition. F10: PROJECT.md Next action and README status now point at M4 pending freeze instead of contradicting HANDOFF in the same tree. F13: name all three zero-gating mechanisms, not one — answers something else, answers correctly behind a modifier the opening-word rule refuses ('Peking duck', 'Electric guitar', 'Insect Beetle'), or misses on morphology. The upward-bias conclusion is unchanged. F14: the 0.5B ceiling (35/41, Wilson lower 0.716) is far above the bar and is not a reason to expect failure; the reasons are M3's in-statistic 19/28, the 0/6 proxy and M1's 33/69 control cell. Removes a "so" that contradicted the preceding sentence. F15: re-wrap the paragraphs the earlier fix commits broke. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- HANDOFF.md | 16 ++++-- PROJECT.md | 12 ++-- README.md | 10 +++- docs/M4-BRIEF.md | 143 +++++++++++++++++++++++++++++++---------------- 4 files changed, 119 insertions(+), 62 deletions(-) diff --git a/HANDOFF.md b/HANDOFF.md index b92463b..ed94b06 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -93,12 +93,16 @@ at 1.5B AND 3B — a *new*, deliberately lenient, uncalibrated constant on the carried from M3's per-cell floor). Realized ns are known from the recorded gated sets (gate arm 41 / 71 / 84); 486 / 844 / 993 of the pool's 492 / 852 / 1,008 cells are genuinely new, the rest are M1-recorded and cap the gate arm at -35/41, 69/71, 82/84 before any forward pass. Design facts found while drafting: four probe clues mention a prime's -spelling (October→september-2, silver→flute-1, China→jade-1, October→opal-2 — -D21 decides their treatment); all 60 roster words pass the D18 span bar on all -three tokenizers (checked in advance); the recorded same-category proxy for -the new pool reads 6/8 (1.5B) and 13/15 (3B) but **0/6 at 0.5B**, so the -off-gate 0.5B floor may genuinely fail — reportable under the standing frame. +35/41, 69/71, 82/84 before any forward pass. Design facts found while +drafting: five probe clues mention a prime's spelling under D5's own prefix +rule (October→september-2, silver→flute-1, China→jade-1, October→opal-2, +Egypt→beetle-2 — the last ungated on all three subjects; D21 decides their +treatment); D9(b)'s owned span-truncation residual re-enters a gate-bearing +arm for the first time since M1 (0 / 2 / 2 cells, carried by a pre-registered +fail-in-place conservative read); all 60 roster words pass the D18 span bar on +all three tokenizers (checked in advance); and M3's own 0.5B subset already +fails M4's bar in-statistic (19/28, Wilson lower 0.4934), so the off-gate 0.5B +floor may genuinely fail — reportable under the standing frame. No runner code exists yet; code only after freeze, cut from `m3_matrix.py`. Standing constraints unchanged: certified environment = `mps` + torch 2.13.0 + diff --git a/PROJECT.md b/PROJECT.md index fd9fead..74c80c9 100644 --- a/PROJECT.md +++ b/PROJECT.md @@ -12,10 +12,14 @@ leaves the other eleven almost untouched (1.5B: diagonal 0/34 vs off-diagonal bit-for-bit on every later run. v1 = M0–M3 per `docs/KICKOFF.md`; S1/S2 stretches optional. -**Next action:** none pre-committed — the v1 chain is closed. Open options: the -S1 stretch (7B lens fit + matrix-lite, ≤ $15 rented GPU), the S2 stretch -(lexical vs semantic scope, which owes the `oracle._BOUNDARY` boundary-class -decision before freezing any non-ASCII list), a write-up, or `/seed-hunt`. +**Next action:** close-out stage **M4, the vocabulary collateral strip** — +decided by Kyle 2026-07-28 over closing immediately and over the stretches. +`docs/M4-BRIEF.md` is written and in adversarial review (PR #10); decisions +**D19–D22 await Kyle's freeze**, and no runner code exists until they do. After +M4: write-up + `/seed-hunt`. The S1 (7B) and S2 (lexical vs semantic scope) +stretches were declined for this repo and banked as idea #13 in +`~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md`; they compete in the +seed-hunt on equal terms. **Key facts** - Fact — Anchor: S4b (dim-stage), concept-specific off-switch at 1.5B, +.727 diff --git a/README.md b/README.md index 818e4c3..3a7f6fb 100644 --- a/README.md +++ b/README.md @@ -70,9 +70,13 @@ deterministic prefix rule on the recorded 3-token span (decision D9b, frozen before any run), M1's published numbers stand untouched, and the re-score is published beside them as a labelled reanalysis (D10a) in which the contrast survives on every subject and the dark categories light up. **The v1 chain -(M0–M3) is now closed; the S1 (7B) and S2 (lexical vs semantic scope) stretches -remain optional.** Models: Qwen2.5-0.5B/1.5B/3B-Instruct, local MPS, -forward-only; core chain $0. +(M0–M3) is now closed.** In progress: close-out stage **M4, the vocabulary +collateral strip** (12 characterized directions × all 60 concepts), which +measures the one thing M3's near-white grid does *not* show — that deleting +France spares the other 48 concepts. Its brief is in review and its decisions +are not yet frozen. The S1 (7B) and S2 (lexical vs semantic scope) stretches +were declined for this repo and banked for a future seed-hunt. Models: +Qwen2.5-0.5B/1.5B/3B-Instruct, local MPS, forward-only; core chain $0. Full brief: [`docs/KICKOFF.md`](docs/KICKOFF.md). The 12-idea backlog this was picked from: dim-stage diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md index 9f00313..9739133 100644 --- a/docs/M4-BRIEF.md +++ b/docs/M4-BRIEF.md @@ -41,10 +41,9 @@ descriptive finding. Finding 1 said collateral concentrates on a few fragile probe columns are a near-fresh sample — 6 / 8 / 15 of their gate-bearing cells are M1-recorded, the other 486 / 844 / 993 are new: either near-total sparing holds and the fragile-probe story stays a curiosity of the 12, or more -silver-like columns exist among the 48 — -both outcomes are findings, and a *failed* gate (a low survival floor across -the wider vocabulary) would be the pre-committed reportable headline against -the specificity story's reach. +silver-like columns exist among the 48 — both outcomes are findings, and a +*failed* gate (a low survival floor across the wider vocabulary) would be the +pre-committed reportable headline against the specificity story's reach. Because the strip contains, cell-for-cell, almost everything M1 and M3 already recorded on these conditions, it also carries the standing re-certification a @@ -87,29 +86,42 @@ recorded artifacts and the frozen `oracle.py` — no new model runs):** 116** (0.5B / 1.5B / 3B) of 180 items. Non-subset gated items — the gate arm's n: **41 / 71 / 84**. Non-subset concepts with ≥ 1 gated item: **23 / 41 / 43** of 48 (zero-gated: 25 / 7 / 5). Those concepts drop out because - **the model names something else**, not because of token geometry: under + **the model answers something else, or answers correctly in a form D9(b)'s + opening-word rule refuses** — never because of token geometry: under D9(b) the gate is a prefix on the 3-token span, so it is insensitive to token count up to 3, and four of the seven 1.5B zero-gated concepts are single-token bare. The recorded clean spans say it plainly — `saturn-1/2/3 → 'Jupiter'`, `neptune-1 → 'Jupiter'`, `bronze-2 → 'Aluminum'`, - `monday-1/2/3 → 'Friday'`, `thursday-1/3 → 'Friday'`, `moth-1 → 'Wasp'`; - `ant-1/2 → 'Ants'` is a morphology miss (a plural closes no boundary). - So the gate arm is **the concepts this subject already names correctly** — a - *competence selection*, and confidently-named concepts are plausibly the - robust ones, which biases the measured sparing floor **upward**. That is the + `monday-1/2/3 → 'Friday'`, `thursday-1/3 → 'Friday'`, `moth-1 → 'Wasp'`. + Two further mechanisms sit beside that one and are named rather than folded + in: `ant-1/2 → 'Ants'` is a **morphology** miss (a plural closes no + boundary), and D9(b)'s opening-word rule refuses **correct answers behind a + modifier** — `duck-3 → 'Peking duck'` and `beetle-3 → 'Insect Beetle'` at 3B, + `eagle-1 → 'American eagle'`, `guitar-2 → 'Electric guitar'` and + `bear-3 → 'Polar bear'` at 0.5B. All three mechanisms point the same way: + the gate arm is **the concepts this subject already names, in the form the + oracle accepts** — a *competence selection*, and confidently-named concepts + are plausibly the robust ones, which biases the measured sparing floor + **upward**. That is the same enrichment mechanism as F12's selection-enrichment finding, now on the probe side; it is why the claim is sparing across the *measurable* vocabulary, said exactly that way. -- **Four probe clues mention a prime's spelling — a confound M3's design never +- **Five probe clues mention a prime's spelling — a confound M3's design never had.** M1's leak guard (D5) bars a clue from leaking its *own* concept or control; M3 verified no cross-mentions *within the 12*. Widening probes to - 180 items surfaces exactly four (prime, item) pairs where the item's clue - contains the deleted word at a word boundary: **October→september-2, - silver→flute-1, China→jade-1, October→opal-2**. In those cells a miss cannot - distinguish "collateral damage to naming machinery" from "the clue's own - text lost a word it references." All four items gate at both gate-bearing - subjects (only jade-1 gates at 0.5B), so the cells are live. D21 decides - their treatment. + 180 items surfaces five (prime, item) pairs, scanned with **D5's own rule** — + no word of the clue may *start with* the string, case-insensitive, plus that + string's `forbidden_forms` entries — not the narrower whole-word match: + **October→september-2, silver→flute-1, China→jade-1, October→opal-2**, and + **Egypt→beetle-2** ("Ancient **Egyptians** carved amulets of the scarab, one + kind of this insect"), which a whole-word scan misses because the clue + inflects the prime. In those cells a miss cannot distinguish "collateral + damage to naming machinery" from "the clue's own text lost a word it + references." The first four items gate at both gate-bearing subjects (only + jade-1 gates at 0.5B), so those cells are live; `beetle-2` is **ungated on + all three subjects** (clean span 'Scarab' / 'scarab'), so it contributes no + gate-bearing cell today and is listed so a future re-gate cannot silently + miss it. D21 decides their treatment. - **D9(b)'s owned span-truncation residual re-enters a gate-bearing arm for the first time since M1.** `oracle.py`'s frozen wording owns a residual: for the three concepts whose bare spelling *fills* the 3-token span — beetle, @@ -158,15 +170,24 @@ recorded artifacts and the frozen `oracle.py` — no new model runs):** 82/84** at 0.5B / 1.5B / 3B (0.5B: all 6 proxy cells miss; 1.5B: `july-3` and `venus-3` miss; 3B: `guitar-2` and `neptune-1` miss). At 1.5B the bar needs 44 of 71, so the ceiling leaves real room; at 3B likewise. Where the - proxy is genuinely predictive is the *rest* of the pool. The pool itself is - mostly cross-category, so a - high floor is the prediction at the gate-bearing subjects. At 0.5B the same - proxy reads **0/6** — consistent with M1's full-battery 0.5B control cell - (33/69) and with F12's selection-enrichment finding (the subset-12's 0.5B - robustness came partly from S1's selection rule), and *not* with M3's clean - 0.5B subset floor. The strip may well fail its floor at 0.5B — off-gate, - under the standing any-direction-damage frame, and reportable as the - measured resolution of that tension. + proxy is genuinely predictive is the *rest* of the pool, which is mostly + cross-category — so a high floor is the prediction at the gate-bearing + subjects. +- **The 0.5B prediction, read in M4's own statistic.** The strongest and most + directly comparable evidence is not the 6-cell proxy: it is M3's recorded + subset, recomputed under **M4's gate statistic**. M3's per-item + survives-all-11 off-diagonal reads **19/28 = 0.679 at 0.5B, Wilson lower + bound 0.4934** — *below 0.5*, i.e. **M3's own 0.5B subset would already fail + M4's bar**. The same field passes at both gate-bearing subjects (1.5B 29/34, + lower 0.6987; 3B 27/32, lower 0.6825). So there is no "tension" between M3's + clean 0.5B floor and a possible M4 failure to resolve — M3's 0.5B floor was + clean only under the *cluster-mean per-cell* statistic, and switching to the + conjunction flips it. That is also the sharpest illustration of why the 0.5 + constant carries no meaning across statistics (D20). The 0.5B strip may well + fail its floor — off-gate, under the standing any-direction-damage frame, + and consistent in advance with M3's own numbers, M1's full-battery 0.5B + control cell (33/69), the 0/6 proxy, and F12's selection-enrichment finding + (the subset-12's 0.5B robustness came partly from S1's selection rule). ## Dispositions for the two inherited obligations (explicit, per HANDOFF) @@ -237,10 +258,13 @@ JSON):** > conjunction**, not on per-cell survival, and is **not** M3's per-cell > floor: under independence it corresponds to a per-cell survival of > 0.5^(1/12) ≈ **0.944**, and its stringency depends on the deletion count - > (12) as much as on the sparing rate. The M4 verdict is the AND over 1.5B - > and 3B; 0.5B runs and is reported under its standing - > any-direction-damage frame, never gate-bearing. Gate-arm n (41 / 71 / 84) - > < MIN_N = 20 ⇒ pre-declared UNDERPOWERED and no claim. + > (12) as much as on the sparing rate. The bar is read **only when the 468 + > M3-recorded and 255 M1-recorded cells in the strip reproduce their recorded + > outcomes bit-for-bit**; any mismatch is INVALID and there is no verdict. + > The M4 verdict is the AND over 1.5B and 3B; 0.5B runs and is reported under + > its standing any-direction-damage frame, never gate-bearing. Gate-arm + > n < MIN_N = 20 ⇒ pre-declared UNDERPOWERED and no claim (realized n = + > 41 / 71 / 84 from the recorded gated sets). *Why this shape.* The strip's question is a **level** question — "is the floor high?" — not an ordering question; M3 already settled the ordering. @@ -275,6 +299,17 @@ JSON):** as M3's floor. It is a floor bar, not an effect-size claim; the descriptive numbers carry the actual size. + *Why the re-certification clause is inside the wording.* Every prior stage's + gate compared an intervened arm against another **measured** arm, so a dead + intervention could never pass one. M4's bar is single-clause and reads only + the off-target survival rate — so read in isolation, an ablation that did + nothing at all would score ~100% survival and print VOCAB-SPARING. In + practice the strip's re-run of M3's 432 matrix cells and M1's 36 + `primed_late` cells catches exactly that, and any mismatch exits INVALID — + but that guarantee lived in D19's design and the runner's exit code, not in + the sentence a write-up quotes. Putting it in `GATE_WORDING` means the + sentence cannot be quoted out of its own precondition. + *Trade-off, owned:* survives-all-12 is the strictest sparing statistic; a single fragile cell fails an item, and the correlation structure across the 12 deletions (unmeasured until this run) decides how harsh that is. That @@ -346,7 +381,7 @@ cluster-mean per-cell floor on the new pool (M3's F15 readout, reference line **row profiles** (does any of the 12 damage the wider vocabulary?) and per-probe **column profiles** over all 60 concepts (finding 1's mostly out-of-sample test — are there silver-like columns among the 48?); the within- vs -cross-category split of the new pool; the four confound cells (D21); mean +cross-category split of the new pool; the five confound pairs (D21); mean concept mass per cell under D13's standing scope; the 0.5B floor under its standing frame (the recorded proxy predicts it may fail — that outcome is a finding, not a failure). @@ -357,22 +392,28 @@ INVALID before the checkpoint loads (the D18-companion shape, carried from `m3_matrix.py`); `--dry-run` validates and stops; `--limit` is smoke, never a result; M4 refuses M1 or M3 artifacts that were themselves not results. -### D21 — The four cross-mention cells +### D21 — The five cross-mention cells + +*The list freezes here, scanned with D5's own prefix + `forbidden_forms` rule: +**October→september-2, silver→flute-1, China→jade-1, October→opal-2, +Egypt→beetle-2**. Four are gated today (one at 0.5B); `beetle-2` is ungated on +all three subjects, so it carries no gate-bearing cell now but is listed so a +future re-gate cannot silently acquire one.* - **(a) Keep them in the gate-bearing pool; report them as a named confound - row (recommended).** The four cells stay in every pooled arm and in their - items' survives-all-12 conjunctions, and the results section reports each - cell's outcome individually under its named confound. *Why:* a confounded - miss can only **lower** the floor — the bias runs against the gate, the one - direction this project ships owned. Excluding them would delete only cells - that could hurt the claim, which is the anti-conservative move the lineage - never makes. Four cells of 852 (1.5B) cannot carry a verdict either way; - what they can do is mislead a *reader* of the column profiles, and the named - row prevents that. + row (recommended).** The cells stay in every pooled arm and in their items' + survives-all-12 conjunctions, and the results section reports each cell's + outcome individually under its named confound. *Why:* a confounded miss can + only **lower** the floor — the bias runs against the gate, the one direction + this project ships owned. Excluding them would delete only cells that could + hurt the claim, which is the anti-conservative move the lineage never makes. + Four cells of 852 (1.5B) cannot carry a verdict either way; what they can do + is mislead a *reader* of the column profiles, and the named row prevents + that. - **(b) Pre-registered exclusion from gate-bearing pools, reported as texture.** Cleaner causal story per cell, but it is evidence-removal in the gate's favour — rejected on the standing bias rule. Not recommended. -- **(c) Drop the four items entirely.** Loses their clean cells and their 11 +- **(c) Drop the five items entirely.** Loses their clean cells and their unconfounded prime cells for no reason. Not recommended. ### D22 — Run-time instrument bars (D18 carried, widened to the full roster) @@ -399,10 +440,10 @@ result; M4 refuses M1 or M3 artifacts that were themselves not results. | M4 exists at all (post-KICKOFF stage) | KICKOFF's frozen M0–M3 chain | Kyle-picked close-out (2026-07-28) that closes M3's stated bound before write-up; KICKOFF's scope decisions unrelitigated; S1/S2 declined and banked (idea #13) | | Level-bar gate (Wilson lower bound vs a constant) | the lineage's Newcombe ordering gates | The ordering is M3's settled result; the strip's question is a level question | | A **new, uncalibrated, sole-dispositive** 0.5 constant | D17's 0.5, which was per-cell and never dispositive | The statistic changed (12-fold conjunction, ≈ 0.944 per-cell under independence) and the status changed (qualifier → gate), so no provenance transfers; owned as deliberately lenient, pre-registered before any new cell, fitted to none, with the per-cell equivalence frozen into `GATE_WORDING` | -| Four cross-mention (prime, item) cells kept in gate-bearing pools | M3's verified no-cross-mention property | M1's leak guard only bars own-concept/control leaks; the confound biases against the gate; named per-cell reporting (D21a) | +| Five cross-mention (prime, item) pairs kept in gate-bearing pools | M3's verified no-cross-mention property | M1's leak guard only bars own-concept/control leaks; scanned with D5's own prefix + `forbidden_forms` rule (a whole-word scan misses `Egypt→beetle-2`, which the clue inflects); the confound biases against the gate; named per-cell reporting (D21a) | | `oracle.py` byte-shared by a fourth consumer (`m4_strip.py`) | cut-from-predecessor rule | Same D9 rationale as the first three consumers: the rule's purpose is byte-identity; pinned by the existing shared-oracle test pattern | | D9(b)'s owned span-truncation residual sits in a gate-bearing arm | `oracle.py`'s scope sentence, "None of the three concepts is in M2's subset" | M4 scores all 60 probes, so beetle / butterfly / trumpet are gated again for the first time since M1 (0 / 2 / 2 gate-arm cells whose span equals the spelling with no boundary observed — `oracle.py`'s own named set); the bias runs *toward* the gate, so it is disclosed per subject and carried by the pre-registered residual-conservative recomputation, fail-in-place (D20), never by editing `oracle.py` | -| Probe-side reach is still the D9(b)-visible roster | "the vocabulary" | 25 / 7 / 5 concepts gate zero items — a **competence selection** (the model names something else; the gate is token-count-insensitive up to 3), not tokenizer geometry; that selection plausibly enriches for robust concepts and biases the floor **upward** (F12's enrichment mechanism, probe side); the claim is sparing across the *measurable* vocabulary, said exactly that way | +| Probe-side reach is still the D9(b)-visible roster | "the vocabulary" | 25 / 7 / 5 concepts gate zero items — a **competence selection** (the model answers something else, answers correctly behind a modifier the opening-word rule refuses, or misses on morphology; the gate is token-count-insensitive up to 3), not tokenizer geometry; that selection plausibly enriches for robust concepts and biases the floor **upward** (F12's enrichment mechanism, probe side); the claim is sparing across the *measurable* vocabulary, said exactly that way | Standing owned rows carry unchanged: naming-only gate (K2), lens provenance (K3), S2-stratum space-keyed directions (D11), mass-channel scope (D13), @@ -422,7 +463,7 @@ now (a run that disagrees is an INVALID cross-check, not a power surprise): | — of them already recorded in M1 (outcome fixed pre-run) | 6 | 8 | 15 | ceilings 35/41, 69/71, 82/84 | | Concept-level collapse (non-subset concepts ≥ 1 gated item) | 23 | 41 | 43 | yes, all | | Subset diagonal (recorded; M3's cells re-run) | 28 | 34 | 32 | yes | -| Confound cells in the pool (D21a) | 1 | 4 | 4 | named texture | +| Cross-mention confound cells in the pool (D21a; 5 pairs, 4 currently gated) | 1 | 4 | 4 | named texture | | D9(b) residual items in the gate arm (clean arm; span = spelling, no boundary) | 0 | 2 | 2 | named texture | Honesty rows, carried and extended: probe-side clustering (3 items share a @@ -438,9 +479,13 @@ pre-registered ceiling of **69 of 71** — 8 of the 852 cells are M1-recorded an two of them (`july-3`, `venus-3`) already miss — there is real room, and M3's per-item collapse texture (29/34 survived all 11 at 1.5B) points into it; the ceiling is a fact about the arm, not evidence for the gate. At 0.5B the same -recorded cells put the ceiling at **35 of 41** (all 6 proxy cells miss), so the -off-gate floor may well fail there, which would be the first measured -divergence between the subset's 0.5B robustness and the wider roster's. +recorded cells put the ceiling at **35 of 41** (all 6 proxy cells miss) — +`wilson(35, 41)` lower bound 0.716, comfortably above the bar, so that ceiling +is *not* a reason to expect failure. The reasons to expect failure at 0.5B are +the ones stated in the instrument facts: M3's own 0.5B subset already fails +this bar in-statistic (19/28, lower 0.4934), the 0/6 proxy rate, and M1's 33/69 +0.5B control cell. A 0.5B failure would be the first measured divergence +between the subset's 0.5B robustness and the wider roster's. ## Wall-clock plan From edb53870ab08717ddd372d728446685dd59e6a41 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Tue, 28 Jul 2026 22:04:09 -0500 Subject: [PATCH 05/13] =?UTF-8?q?docs(m4):=20review=20round=203=20fixes=20?= =?UTF-8?q?=E2=80=94=20F16,=20F17,=20F18,=20F15=20re-fix?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit F16 [should-fix]: the F11 fix swapped a selector matching EVERY cell for one matching NONE. "Equals the scored concept's spelling exactly" is case-exact; recorded residual spans are 'Beetle' / 'Butterfly' / 'Trumpet' while roster spellings are lowercase, so a verbatim runner would select zero cells and turn the sole mitigation for a bias-toward-the-gate into a no-op. Verified: over all 540 cells x 3 subjects, case-exact selects 0; case-insensitive selects exactly the pre-registered 0 / 2 / 2. The selector now states case-insensitive comparison "exactly as oracle.says_concept_prefix compares (re.IGNORECASE)", and says why the case rule decides the set rather than leaving it as a detail. F17: the F10 fix rewrote each file's forward-looking prose but not the status lines above it, leaving PROJECT.md and README.md self-contradictory (each was internally consistent before). PROJECT.md's Status paragraph and README's Status line now both carry M4-in-flight, and README's milestone table gains an M4 row with S1/S2 relabelled banked. F18: "exactly the set oracle.py's docstring names" was four cells against the oracle's stated six — the docstring enumerates across arms (butterfly-1 x2 arms: clean and control_late). Restated as the gate-arm members of that set, with the arm/subject scoping spelled out. F15 re-fix: the :383 wrap was missed and the F13 fix introduced a new orphan at :105. Both repaired; full re-scan of all four files is clean. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- PROJECT.md | 6 ++++-- README.md | 8 +++++--- docs/M4-BRIEF.md | 48 +++++++++++++++++++++++++++++------------------- 3 files changed, 38 insertions(+), 24 deletions(-) diff --git a/PROJECT.md b/PROJECT.md index 74c80c9..2d17bda 100644 --- a/PROJECT.md +++ b/PROJECT.md @@ -9,8 +9,10 @@ deleting one concept's direction at the late third silences that concept and leaves the other eleven almost untouched (1.5B: diagonal 0/34 vs off-diagonal 363/374). M2 PASSED (LATE-LOCALIZED at 1.5B and 3B), M1 PASSED (BREADTH-SPECIFIC at 1.5B and 3B), M0 PASSED (2026-07-27); all re-certified -bit-for-bit on every later run. v1 = M0–M3 per `docs/KICKOFF.md`; S1/S2 -stretches optional. +bit-for-bit on every later run. v1 = M0–M3 per `docs/KICKOFF.md`. **Close-out +stage M4 (the vocabulary collateral strip) is now in flight** — brief written +and in review, decisions D19–D22 not yet frozen, no runner code yet. The S1/S2 +stretches were declined for this repo and banked (idea #13). **Next action:** close-out stage **M4, the vocabulary collateral strip** — decided by Kyle 2026-07-28 over closing immediately and over the stretches. diff --git a/README.md b/README.md index 3a7f6fb..11038fd 100644 --- a/README.md +++ b/README.md @@ -21,13 +21,15 @@ pre-registered gates frozen as code before any run): | M1 | Breadth | How much of the (measurable) vocabulary has an off-switch? | | M2 | Localization + dose | Where does the switch live, and how much removal does it take? | | M3 | Specificity | The full prime × probe collateral matrix. | -| S1 (stretch) | Scale | The specificity-emergence curve, extended to 7B. | -| S2 (stretch) | Scope | Token mute button or concept mute button? | +| M4 (close-out) | Vocabulary collateral | Does deleting one concept spare the *other 48*? | +| S1 (banked) | Scale | The specificity-emergence curve, extended to 7B. | +| S2 (banked) | Scope | Token mute button or concept mute button? | **The honest framing:** an effect *found during a replication, characterized here* — the anchor is dim-stage's own recorded result, not a paper claim. -**Status: M3 PASSED 2026-07-28 — the v1 chain is complete.** On the full 12 × 12 +**Status: M3 PASSED 2026-07-28 — the v1 chain is complete; close-out stage M4 +is in flight (brief in review, decisions not yet frozen).** On the full 12 × 12 prime × probe matrix at the switch's home band, deleting one concept's direction silences that concept and leaves the other eleven almost untouched: at 1.5B the diagonal names **0/34** while the pooled off-diagonal names **363/374** diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md index 9739133..32fffa9 100644 --- a/docs/M4-BRIEF.md +++ b/docs/M4-BRIEF.md @@ -102,10 +102,9 @@ recorded artifacts and the frozen `oracle.py` — no new model runs):** the gate arm is **the concepts this subject already names, in the form the oracle accepts** — a *competence selection*, and confidently-named concepts are plausibly the robust ones, which biases the measured sparing floor - **upward**. That is the - same enrichment mechanism as F12's selection-enrichment finding, now on the - probe side; it is why the claim is sparing across the *measurable* - vocabulary, said exactly that way. + **upward**. That is the same enrichment mechanism as F12's + selection-enrichment finding, now on the probe side; it is why the claim is + sparing across the *measurable* vocabulary, said exactly that way. - **Five probe clues mention a prime's spelling — a confound M3's design never had.** M1's leak guard (D5) bars a clue from leaking its *own* concept or control; M3 verified no cross-mentions *within the 12*. Widening probes to @@ -133,14 +132,20 @@ recorded artifacts and the frozen `oracle.py` — no new model runs):** probe side to all 60, so their items re-enter a **gate-bearing** pool. The residual condition is exact and narrower than "a 3-token word": the recorded span, after stripping leading whitespace, **equals the concept's - spelling with nothing following it**, so no boundary character is observed. + spelling with nothing following it**, compared case-insensitively as + `oracle.says_concept_prefix` compares, so no boundary character is observed. ("Fills the 3-token span" would not distinguish anything — every recorded - `greedy_3` is exactly 3 tokens by construction.) On that reading the - gate-arm residual cells are **0 / 2 / 2** at 0.5B / 1.5B / 3B — 1.5B - `beetle-1` ('Beetle') and `butterfly-1` ('Butterfly') of 71; 3B `trumpet-3` - ('Trumpet') and `butterfly-1` ('Butterfly') of 84; none gated at 0.5B. This - is exactly the set `oracle.py`'s frozen docstring already names ("beetle-1, - butterfly-1 ×2 arms, trumpet-3"). The other trumpet cells are *not* + `greedy_3` is exactly 3 tokens by construction; and a case-*exact* + comparison would select nothing, since the spans are capitalised and the + roster spellings are not.) On that reading the gate-arm residual cells are + **0 / 2 / 2** at 0.5B / 1.5B / 3B — 1.5B `beetle-1` ('Beetle') and + `butterfly-1` ('Butterfly') of 71; 3B `trumpet-3` ('Trumpet') and + `butterfly-1` ('Butterfly') of 84; none gated at 0.5B. These are the + gate-arm members of the same residual set `oracle.py`'s frozen docstring + names — the docstring counts **six recorded M1 cells** ("beetle-1, + butterfly-1 ×2 arms, trumpet-3"), enumerating across arms (`butterfly-1` + in both `clean` and `control_late`), where the four above are clean-arm + cells on the two gate-bearing subjects. The other trumpet cells are *not* residual: `trumpet-1` at 1.5B and `trumpet-1` / `trumpet-2` at 3B record `'Trumpet<|im_end|>'`, and `<|im_end|>` closes a word under `oracle._BOUNDARY` — capitalisation is why, since generated `Trumpet` is two @@ -354,9 +359,14 @@ Beside the gate as scored, the same gate statistic is recomputed with every **residual cell** re-scored as a **miss** — the maximally conservative reading of the boundary D9(b) cannot observe. A residual cell is one whose recorded span, after stripping leading whitespace, **equals the scored concept's -spelling exactly, with nothing following it** (no boundary character observed). -That is the selector, stated once here and implemented verbatim: it is *not* -"the span fills the 3-token window", which every recorded cell does. In the +spelling with nothing following it — compared case-insensitively, exactly as +`oracle.says_concept_prefix` compares (`re.IGNORECASE`)**, so that no boundary +character is observed. The case rule is stated because it decides the set, not +as a detail: the recorded spans are `'Beetle'`, `'Butterfly'`, `'Trumpet'` +while the roster spellings are lowercase, so a case-*exact* comparison would +select **zero** cells and silently turn this mitigation into a no-op. That is +the selector, stated once here and implemented verbatim: it is *not* "the span +fills the 3-token window", which every recorded cell does. In the recorded clean arm the residual cells are the 0 / 2 / 2 gate-arm items named in the instrument facts; in the ablated cells the set is whatever the run records, computed from the same recorded spans by the same selector. @@ -380,11 +390,11 @@ cluster-mean per-cell floor on the new pool (M3's F15 readout, reference line 0.5, the M3-comparable view); the ordering contrast from option (b); per-prime **row profiles** (does any of the 12 damage the wider vocabulary?) and per-probe **column profiles** over all 60 concepts (finding 1's mostly -out-of-sample test — are there silver-like columns among the 48?); the within- vs -cross-category split of the new pool; the five confound pairs (D21); mean -concept mass per cell under D13's standing scope; the 0.5B floor under its -standing frame (the recorded proxy predicts it may fail — that outcome is a -finding, not a failure). +out-of-sample test — are there silver-like columns among the 48?); the +within- vs cross-category split of the new pool; the five confound pairs +(D21); mean concept mass per cell under D13's standing scope; the 0.5B floor +under its standing frame (M3's own subset already fails this bar in-statistic, +so a 0.5B failure is a finding, not a failure of the run). **Verdict precedence, frozen** in `strip_verdict()` (`m4_verdict.py`): NOT A RESULT > DEGENERATE > UNDERPOWERED > the level bar. Wrong-arm inputs exit From 90b994c21c3b2905824abf121f3062d2d2b48c42 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Wed, 29 Jul 2026 01:10:24 -0500 Subject: [PATCH 06/13] =?UTF-8?q?docs(m4):=20freeze=20D19-D22=20=E2=80=94?= =?UTF-8?q?=20(a)=20across=20the=20board,=20F8=20resolved?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Kyle froze all four decisions 2026-07-29 at the review gate: D19(a) full 12 x 180 strip + clean re-run (2,340 cells/subject), with the embedded 255-cell M1 and 468-cell M3 re-certifications graded first. D20(a) the single-clause survives-all-12 level gate as written. D21(a) all five cross-mention pairs kept in the gate-bearing pool with named per-cell reporting. D22(a) both run-time instrument bars over all 60 scored words + the 12 direction words. F8 (the verdict-label choice, deliberately left open as a design question rather than fixed as a defect) resolved: the label stays VOCAB-SPARING and the realized survival proportion rides INSIDE the verdict string — "VOCAB-SPARING (k/n survive all 12 = ; Wilson 95% lower )". This is M3's ON A DAMAGED FLOOR pattern applied to a level bar: the label is what a write-up quotes and it is frozen into every results JSON, so the number that qualifies it travels with it rather than living in prose. The brief now states why — at the 1.5B pass point 27 of 71 gated items (38%) are damaged by at least one deletion, and at the nominal 0.5 it would be half. Freeze note added to the brief per the M3 precedent; full DECISIONS.md entries D19-D22 land with the M4 code PR, per the M0-M3 pattern. PROJECT.md next action moves to building the runner; HANDOFF.md records the freeze. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- HANDOFF.md | 14 +++++++++++--- PROJECT.md | 14 +++++++------- docs/M4-BRIEF.md | 39 ++++++++++++++++++++++++++++++++++++++- 3 files changed, 56 insertions(+), 11 deletions(-) diff --git a/HANDOFF.md b/HANDOFF.md index ed94b06..01c4bb2 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -84,9 +84,17 @@ vocabulary collateral strip as close-out stage M4, then write-up + idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md` — they compete in the seed-hunt on equal terms, no incumbent's privilege. -**Where M4 stands: `docs/M4-BRIEF.md` is written and in review; decisions -D19–D22 await Kyle's freeze.** The stage: 12 subset primes × all 180 M1 items -(2,340 cells/subject), gate proposed as a single-clause VOCAB-SPARING level +**Where M4 stands: `docs/M4-BRIEF.md` is written, reviewed across three +adversarial rounds, and D19–D22 are FROZEN (Kyle, 2026-07-29) — (a) across the +board.** The verdict label was resolved per review F8: the label stays +`VOCAB-SPARING` and the realized survival proportion rides inside the verdict +string, the M3 `ON A DAMAGED FLOOR` pattern applied to a level bar. Full +`DECISIONS.md` entries D19–D22 land with the M4 code PR, per the M0–M3 +pattern. Round 4 of the review was authorized beyond the three-dispatch cap to +verify the `edb5387` fixes and the freeze commit; no merge until it is clean. + +The stage: 12 subset primes × all 180 M1 items +(2,340 cells/subject), gate = a single-clause VOCAB-SPARING level bar on the non-subset pool (per-item survives-all-12, Wilson lower bound ≥ 0.5 at 1.5B AND 3B — a *new*, deliberately lenient, uncalibrated constant on the 12-fold conjunction, ≈ 0.944 per-cell under independence, explicitly not diff --git a/PROJECT.md b/PROJECT.md index 2d17bda..615158d 100644 --- a/PROJECT.md +++ b/PROJECT.md @@ -11,14 +11,14 @@ leaves the other eleven almost untouched (1.5B: diagonal 0/34 vs off-diagonal (BREADTH-SPECIFIC at 1.5B and 3B), M0 PASSED (2026-07-27); all re-certified bit-for-bit on every later run. v1 = M0–M3 per `docs/KICKOFF.md`. **Close-out stage M4 (the vocabulary collateral strip) is now in flight** — brief written -and in review, decisions D19–D22 not yet frozen, no runner code yet. The S1/S2 -stretches were declined for this repo and banked (idea #13). +and reviewed, **decisions D19–D22 frozen 2026-07-29**, no runner code yet. The +S1/S2 stretches were declined for this repo and banked (idea #13). -**Next action:** close-out stage **M4, the vocabulary collateral strip** — -decided by Kyle 2026-07-28 over closing immediately and over the stretches. -`docs/M4-BRIEF.md` is written and in adversarial review (PR #10); decisions -**D19–D22 await Kyle's freeze**, and no runner code exists until they do. After -M4: write-up + `/seed-hunt`. The S1 (7B) and S2 (lexical vs semantic scope) +**Next action:** build the M4 runner. `docs/M4-BRIEF.md` is written and +adversarially reviewed (PR #10); **D19–D22 were frozen 2026-07-29 — (a) across +the board** — so the next step is code: `m4_strip.py` cut from `m3_matrix.py`, +`m4_verdict.py`, `test_m4.py`, and the D19–D22 entries appended to +`docs/DECISIONS.md` in that same code PR. After M4: write-up + `/seed-hunt`. The S1 (7B) and S2 (lexical vs semantic scope) stretches were declined for this repo and banked as idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md`; they compete in the seed-hunt on equal terms. diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md index 32fffa9..b69f5c6 100644 --- a/docs/M4-BRIEF.md +++ b/docs/M4-BRIEF.md @@ -217,6 +217,26 @@ list before the wording freezes. ## Decisions to freeze (Kyle picks; recommendations flagged) +*Frozen (Kyle, 2026-07-29): **D19 (a)** the full 12 × 180 strip plus a full +clean re-run (2,340 cells/subject) with the embedded 255-cell M1 and 468-cell +M3 re-certifications graded first; **D20 (a)** the single-clause +survives-all-12 level gate as written, with the verdict-label choice resolved +per review F8 — the label stays `VOCAB-SPARING` and the **realized survival +proportion rides inside the verdict string**, the M3 `ON A DAMAGED FLOOR` +pattern applied to a level bar; **D21 (a)** all five cross-mention pairs kept +in the gate-bearing pool with named per-cell reporting; **D22 (a)** both +run-time instrument bars over all 60 scored words plus the 12 direction words. +Full `DECISIONS.md` entries (D19–D22) land with the M4 code PR, per the +M0/M1/M2/M3 pattern. Amended pre-freeze at PR #10's adversarial review, across +three rounds: F1–F4 (should-fix) fixed and verified; F11–F12 (should-fix, +defects of the F2 fix) fixed and verified; F16 (should-fix, a defect of the +F11 fix — a case-exact selector that would have silently matched zero cells) +fixed with F15/F17/F18 at `edb5387`; and all eight deferred nice-to-haves +(F5–F7, F9, F10, F13–F15) pulled in at the freeze on Kyle's call ("pull in all +8"), because F5/F7/F9/F13's text byte-freezes into code once these decisions +are signed. Round 4 was authorized by Kyle beyond the three-dispatch cap to +verify the `edb5387` fixes and this freeze commit.* + ### D19 — Primes × probes: the strip frame (decide first) - **(a) 12 subset primes × all 180 M1 items, plus a full clean re-run @@ -269,7 +289,12 @@ JSON):** > The M4 verdict is the AND over 1.5B and 3B; 0.5B runs and is reported under > its standing any-direction-damage frame, never gate-bearing. Gate-arm > n < MIN_N = 20 ⇒ pre-declared UNDERPOWERED and no claim (realized n = - > 41 / 71 / 84 from the recorded gated sets). + > 41 / 71 / 84 from the recorded gated sets). **Verdict string, + > pre-committed:** the label alone over-reads — clearing a floor bar is + > compatible with a large minority of measurable items damaged — so the + > verdict, whatever it is, carries its **realized survival proportion in the + > same string**, per subject: + > `VOCAB-SPARING (k/n survive all 12 = ; Wilson 95% lower )`. *Why this shape.* The strip's question is a **level** question — "is the floor high?" — not an ordering question; M3 already settled the ordering. @@ -315,6 +340,18 @@ JSON):** the sentence a write-up quotes. Putting it in `GATE_WORDING` means the sentence cannot be quoted out of its own precondition. + *Why the realized proportion rides in the verdict string.* `VOCAB-SPARING` + is the strongest possible phrasing of exactly the over-reading M4 exists to + correct, and the bar it names permits real damage: at the 1.5B pass point + (44/71) **27 of 71** gated non-subset items — 38% — are damaged by at least + one of the 12 deletions, and at the bar's nominal 0.5 it would be half. The + prose owns that, but prose is not what gets quoted; the label is, and the + label is frozen into every results JSON. Rather than rename the verdict away + from the lineage's `M-` shape, M4 carries the number **inside + the verdict string** — the same move M3 made when it attached `ON A DAMAGED + FLOOR` to the verdict rather than leaving the damaged floor in the prose. + The number a write-up quotes then travels with the label it quotes. + *Trade-off, owned:* survives-all-12 is the strictest sparing statistic; a single fragile cell fails an item, and the correlation structure across the 12 deletions (unmeasured until this run) decides how harsh that is. That From 52d381e75f9afdd0a62c2927401fffbd7c73f076 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Wed, 29 Jul 2026 01:10:47 -0500 Subject: [PATCH 07/13] docs(m4): re-wrap two lines broken by the freeze commit Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- HANDOFF.md | 4 ++-- PROJECT.md | 3 ++- 2 files changed, 4 insertions(+), 3 deletions(-) diff --git a/HANDOFF.md b/HANDOFF.md index 01c4bb2..fdac886 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -93,8 +93,8 @@ string, the M3 `ON A DAMAGED FLOOR` pattern applied to a level bar. Full pattern. Round 4 of the review was authorized beyond the three-dispatch cap to verify the `edb5387` fixes and the freeze commit; no merge until it is clean. -The stage: 12 subset primes × all 180 M1 items -(2,340 cells/subject), gate = a single-clause VOCAB-SPARING level +The stage: 12 subset primes × all 180 M1 items (2,340 cells/subject), gate = +a single-clause VOCAB-SPARING level bar on the non-subset pool (per-item survives-all-12, Wilson lower bound ≥ 0.5 at 1.5B AND 3B — a *new*, deliberately lenient, uncalibrated constant on the 12-fold conjunction, ≈ 0.944 per-cell under independence, explicitly not diff --git a/PROJECT.md b/PROJECT.md index 615158d..5488b3d 100644 --- a/PROJECT.md +++ b/PROJECT.md @@ -18,7 +18,8 @@ S1/S2 stretches were declined for this repo and banked (idea #13). adversarially reviewed (PR #10); **D19–D22 were frozen 2026-07-29 — (a) across the board** — so the next step is code: `m4_strip.py` cut from `m3_matrix.py`, `m4_verdict.py`, `test_m4.py`, and the D19–D22 entries appended to -`docs/DECISIONS.md` in that same code PR. After M4: write-up + `/seed-hunt`. The S1 (7B) and S2 (lexical vs semantic scope) +`docs/DECISIONS.md` in that same code PR. After M4: write-up + `/seed-hunt`. +The S1 (7B) and S2 (lexical vs semantic scope) stretches were declined for this repo and banked as idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md`; they compete in the seed-hunt on equal terms. From 1c9e8711a929c260196e93f7dd51da694d263b98 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Wed, 29 Jul 2026 01:12:03 -0500 Subject: [PATCH 08/13] docs(m4): re-flow the two paragraphs the freeze commit re-wrapped Line-by-line patching kept shifting the break point; re-flowed both whole paragraphs at 78 columns instead. Also drops a line in HANDOFF that still said "code only after freeze" now that D19-D22 are frozen. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- HANDOFF.md | 38 +++++++++++++++++++------------------- PROJECT.md | 4 ++-- 2 files changed, 21 insertions(+), 21 deletions(-) diff --git a/HANDOFF.md b/HANDOFF.md index fdac886..4eb9984 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -93,25 +93,25 @@ string, the M3 `ON A DAMAGED FLOOR` pattern applied to a level bar. Full pattern. Round 4 of the review was authorized beyond the three-dispatch cap to verify the `edb5387` fixes and the freeze commit; no merge until it is clean. -The stage: 12 subset primes × all 180 M1 items (2,340 cells/subject), gate = -a single-clause VOCAB-SPARING level -bar on the non-subset pool (per-item survives-all-12, Wilson lower bound ≥ 0.5 -at 1.5B AND 3B — a *new*, deliberately lenient, uncalibrated constant on the -12-fold conjunction, ≈ 0.944 per-cell under independence, explicitly not -carried from M3's per-cell floor). Realized ns are known from the recorded -gated sets (gate arm 41 / 71 / 84); 486 / 844 / 993 of the pool's 492 / 852 / -1,008 cells are genuinely new, the rest are M1-recorded and cap the gate arm at -35/41, 69/71, 82/84 before any forward pass. Design facts found while -drafting: five probe clues mention a prime's spelling under D5's own prefix -rule (October→september-2, silver→flute-1, China→jade-1, October→opal-2, -Egypt→beetle-2 — the last ungated on all three subjects; D21 decides their -treatment); D9(b)'s owned span-truncation residual re-enters a gate-bearing -arm for the first time since M1 (0 / 2 / 2 cells, carried by a pre-registered -fail-in-place conservative read); all 60 roster words pass the D18 span bar on -all three tokenizers (checked in advance); and M3's own 0.5B subset already -fails M4's bar in-statistic (19/28, Wilson lower 0.4934), so the off-gate 0.5B -floor may genuinely fail — reportable under the standing frame. -No runner code exists yet; code only after freeze, cut from `m3_matrix.py`. +The stage: 12 subset primes × all 180 M1 items (2,340 cells/subject), gate = a +single-clause VOCAB-SPARING level bar on the non-subset pool (per-item +survives-all-12, Wilson lower bound ≥ 0.5 at 1.5B AND 3B — a *new*, +deliberately lenient, uncalibrated constant on the 12-fold conjunction, ≈ +0.944 per-cell under independence, explicitly not carried from M3's per-cell +floor). Realized ns are known from the recorded gated sets (gate arm 41 / 71 / +84); 486 / 844 / 993 of the pool's 492 / 852 / 1,008 cells are genuinely new, +the rest are M1-recorded and cap the gate arm at 35/41, 69/71, 82/84 before +any forward pass. Design facts found while drafting: five probe clues mention +a prime's spelling under D5's own prefix rule (October→september-2, +silver→flute-1, China→jade-1, October→opal-2, Egypt→beetle-2 — the last +ungated on all three subjects; D21 decides their treatment); D9(b)'s owned +span-truncation residual re-enters a gate-bearing arm for the first time since +M1 (0 / 2 / 2 cells, carried by a pre-registered fail-in-place conservative +read); all 60 roster words pass the D18 span bar on all three tokenizers +(checked in advance); and M3's own 0.5B subset already fails M4's bar +in-statistic (19/28, Wilson lower 0.4934), so the off-gate 0.5B floor may +genuinely fail — reportable under the standing frame. No runner code exists +yet; with D19–D22 now frozen, code is the next step, cut from `m3_matrix.py`. Standing constraints unchanged: certified environment = `mps` + torch 2.13.0 + transformers 5.13.1 (off it: NOT A RESULT); `m0_anchor.py` stays certified and diff --git a/PROJECT.md b/PROJECT.md index 5488b3d..9137215 100644 --- a/PROJECT.md +++ b/PROJECT.md @@ -19,8 +19,8 @@ adversarially reviewed (PR #10); **D19–D22 were frozen 2026-07-29 — (a) acro the board** — so the next step is code: `m4_strip.py` cut from `m3_matrix.py`, `m4_verdict.py`, `test_m4.py`, and the D19–D22 entries appended to `docs/DECISIONS.md` in that same code PR. After M4: write-up + `/seed-hunt`. -The S1 (7B) and S2 (lexical vs semantic scope) -stretches were declined for this repo and banked as idea #13 in +The S1 (7B) and S2 (lexical vs semantic scope) stretches were declined for +this repo and banked as idea #13 in `~/Projects/j-lens-proj-ideas/jlens-followon-backlog.md`; they compete in the seed-hunt on equal terms. From df96498ce67253f11fe5d49dd29bd20075362657 Mon Sep 17 00:00:00 2001 From: ksdisch Date: Wed, 29 Jul 2026 01:22:22 -0500 Subject: [PATCH 09/13] =?UTF-8?q?docs(m4):=20round=204=20fixes=20=E2=80=94?= =?UTF-8?q?=20F19,=20F20,=20F21=20(F20=20amends=20the=20frozen=20wording)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit F19: the freeze commit updated PROJECT.md and HANDOFF.md but not README.md, so the front page still said "decisions not yet frozen" while the other two said FROZEN — in a repo whose whole discipline is freeze-before-code, that answered "were D19-D22 frozen before the run?" with "no". Both README spots corrected. F20 [should-fix, amends the frozen D20 wording]: the verdict-string clause as frozen carried only the AS-SCORED proportion, so the two pre-registered reads that can flip which number is honest — residual-conservative fail-in-place and the concept-level collapse, each closed in the brief with "the ... numbers are the honest ones to quote" — stayed in prose. That is precisely the failure the F8 resolution was adopted to prevent ("prose is not what gets quoted; the label is"). The window is live: at 1.5B the bar needs k >= 44 (wilson(44,71) lower 0.50342), and fail-in-place removes the 2 residual items, so at an as-scored k of 44 or 45 the JSON would have printed VOCAB-SPARING ... lower 0.503 while the brief's own pre-commitment said 42/71 (lower 0.47539) was honest. M3's precedent is a CONDITIONAL qualifier attached by the runner (m3_matrix.py:179); M4 had borrowed the sentence shape and dropped the mechanism. D20 now carries a pre-declared AS-SCORED ONLY qualifier, attached whenever a conservative read's Wilson lower bound is below 0.5 while the as-scored read's is not, naming each such read and its number. Per D17's carried rule the qualifier scopes a claim and can never create or rescue one; the gate, its 0.5 bar, its arm and its re-certification precondition are all unchanged. F21: "the verdict, whatever it is" was followed by a template hard-coding the pass label, and no failing label appeared anywhere in the brief. NOT VOCAB-SPARING is now named in the wording. Recorded in the brief's freeze note as an amendment POST-FREEZE, PRE-RUN — the M3 precedent (its own round-4 F15/F16 amendment) — and FLAGGED FOR KYLE'S RATIFICATION IN THE MERGE BRIEF. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015YkaGZYHwRoGa4e6AsLNLL --- README.md | 41 ++++++++++++++++++++++------------------- docs/M4-BRIEF.md | 46 +++++++++++++++++++++++++++++++++++++++++++--- 2 files changed, 65 insertions(+), 22 deletions(-) diff --git a/README.md b/README.md index 11038fd..aa8638e 100644 --- a/README.md +++ b/README.md @@ -29,21 +29,22 @@ pre-registered gates frozen as code before any run): here* — the anchor is dim-stage's own recorded result, not a paper claim. **Status: M3 PASSED 2026-07-28 — the v1 chain is complete; close-out stage M4 -is in flight (brief in review, decisions not yet frozen).** On the full 12 × 12 -prime × probe matrix at the switch's home band, deleting one concept's direction -silences that concept and leaves the other eleven almost untouched: at 1.5B the -diagonal names **0/34** while the pooled off-diagonal names **363/374** -(+0.971 [+0.867, +0.983]); at 3B 3/32 vs 343/352 (+0.881 [+0.731, +0.943]). The -same contrast restricted to *same-category* pairs — the arm the lineage's single -control actually tested — is +0.950 and +0.891, likewise CI-clean. Of the -matrix's 132 ordered off-diagonal pairs, **126 had never been measured before** -(the other 6 are M1's own country control cells, which this run re-certifies -bit-for-bit before it reads anything new). What the gate did not -ask: collateral concentrates on a few fragile **probes** rather than being caused -by damaging **primes** (`silver`'s direction damages *nothing*, while `silver` -itself is the most fragile probe in the grid — inverting what a single control -cell had suggested), category-block collateral is CI-clean at 0.5B and dissolves -by 1.5B, and the only imperfect mutes anywhere are `Egypt` and `October` at 3B — +is in flight — brief reviewed, decisions D19–D22 frozen 2026-07-29, runner not +yet written.** On the full 12 × 12 prime × probe matrix at the switch's home +band, deleting one concept's direction silences that concept and leaves the +other eleven almost untouched: at 1.5B the diagonal names **0/34** while the +pooled off-diagonal names **363/374** (+0.971 [+0.867, +0.983]); at 3B 3/32 vs +343/352 (+0.881 [+0.731, +0.943]). The same contrast restricted to +*same-category* pairs — the arm the lineage's single control actually tested — +is +0.950 and +0.891, likewise CI-clean. Of the matrix's 132 ordered +off-diagonal pairs, **126 had never been measured before** (the other 6 are +M1's own country control cells, which this run re-certifies bit-for-bit before +it reads anything new). What the gate did not ask: collateral concentrates on +a few fragile **probes** rather than being caused by damaging **primes** +(`silver`'s direction damages *nothing*, while `silver` itself is the most +fragile probe in the grid — inverting what a single control cell had +suggested), category-block collateral is CI-clean at 0.5B and dissolves by +1.5B, and the only imperfect mutes anywhere are `Egypt` and `October` at 3B — the two concepts pre-registered as the leaky-switch stratum. **M2 PASSED 2026-07-28** — on a pre-registered 12-concept subset, the @@ -75,10 +76,12 @@ survives on every subject and the dark categories light up. **The v1 chain (M0–M3) is now closed.** In progress: close-out stage **M4, the vocabulary collateral strip** (12 characterized directions × all 60 concepts), which measures the one thing M3's near-white grid does *not* show — that deleting -France spares the other 48 concepts. Its brief is in review and its decisions -are not yet frozen. The S1 (7B) and S2 (lexical vs semantic scope) stretches -were declined for this repo and banked for a future seed-hunt. Models: -Qwen2.5-0.5B/1.5B/3B-Instruct, local MPS, forward-only; core chain $0. +France spares the other 48 concepts. Its brief is adversarially reviewed and +its decisions (D19–D22) were frozen 2026-07-29, before any runner code exists +— the lineage's freeze-before-code discipline. The S1 (7B) and S2 (lexical vs +semantic scope) stretches were declined for this repo and banked for a future +seed-hunt. Models: Qwen2.5-0.5B/1.5B/3B-Instruct, local MPS, forward-only; +core chain $0. Full brief: [`docs/KICKOFF.md`](docs/KICKOFF.md). The 12-idea backlog this was picked from: dim-stage diff --git a/docs/M4-BRIEF.md b/docs/M4-BRIEF.md index b69f5c6..70c60d8 100644 --- a/docs/M4-BRIEF.md +++ b/docs/M4-BRIEF.md @@ -237,6 +237,20 @@ fixed with F15/F17/F18 at `edb5387`; and all eight deferred nice-to-haves are signed. Round 4 was authorized by Kyle beyond the three-dispatch cap to verify the `edb5387` fixes and this freeze commit.* +*Amended post-freeze, pre-run, at round 4 (F20 + F21) — the M3 precedent for a +review-driven amendment before any cell is run, **flagged for Kyle's +ratification in the merge brief**: the verdict string as first frozen carried +only the **as-scored** proportion, so the two pre-registered reads that can flip +which number is honest (residual-conservative fail-in-place; concept-level +collapse) stayed in prose — reproducing the exact failure F8 was resolved to +prevent. D20's wording now carries the pre-declared **AS-SCORED ONLY** +qualifier, attached conditionally by the runner whenever a conservative read's +Wilson lower bound falls below 0.5 while the as-scored read's does not, and +names the failing label (`NOT VOCAB-SPARING`) that the single pass-label +template had left unstated. The gate, its 0.5 bar, its arm and its precondition +are unchanged — this scopes the claim and, per D17's carried rule, can never +create or rescue one.* + ### D19 — Primes × probes: the strip frame (decide first) - **(a) 12 subset primes × all 180 M1 items, plus a full clean re-run @@ -292,9 +306,21 @@ JSON):** > 41 / 71 / 84 from the recorded gated sets). **Verdict string, > pre-committed:** the label alone over-reads — clearing a floor bar is > compatible with a large minority of measurable items damaged — so the - > verdict, whatever it is, carries its **realized survival proportion in the - > same string**, per subject: - > `VOCAB-SPARING (k/n survive all 12 = ; Wilson 95% lower )`. + > verdict, whichever way it goes, carries its **realized survival proportion + > in the same string**, per subject, using `VOCAB-SPARING` when the bar is + > cleared and `NOT VOCAB-SPARING` when it is not: + > `