Skip to content

Laguna §2: the thinking gate is two axes, not one dose curve (closes #11)#12

Merged
TheTom merged 3 commits into
mainfrom
laguna-two-axis-gate
Jul 26, 2026
Merged

Laguna §2: the thinking gate is two axes, not one dose curve (closes #11)#12
TheTom merged 3 commits into
mainfrom
laguna-two-axis-gate

Conversation

@TheTom

@TheTom TheTom commented Jul 26, 2026

Copy link
Copy Markdown
Owner

Closes #11. @Blackwellboy caught a bad inference in §2 and offered a better version of the curve. Both taken.

The correction

The sentence he quoted had already been removed in 221a311 during the dose-curve retraction, but the substance was still live elsewhere in §2, because I was still presenting the compression numbers (3536 to 745 to 282) as a single suppression progression without saying that firing went up across the last step. So this is not a no-op.

His C7 vs C8 differ by exactly one thing, tool schemas in the request:

condition apparatus fired median think-tok ceiling hits
C7 full agent prompt 24/40 (60%) 745 5
C8 C7 + tool schemas 29/40 (72%) 282 2

Reasoning length down 62%, firing up 12 points, second-highest firing condition in his whole grid. I re-derived every number in this PR from his published grid_turns.jsonl rather than transcribing his table, and all ten conditions match exactly.

§2 now states the two axes separately, with C7/C8 as the worked example of them disagreeing. His point about why it matters is the part I have quoted into the guide in substance: collapsing firing and length into one "suppressor" axis is what made the persona-times-task gate confusing for three people independently.

The curve

Taking his offer. The three-stack composite is gone, replaced by his ten single-stack points (one rev, one prompt set, n=40 each, varying only apparatus). He was right that a curve assembled from three stacks with three different prompt sets is the weakest form of the argument. Our rows and @Defilan's are demoted to corroborating the ordering, with a note that absolute levels are stack-specific, since his bare rate is 75% against our ~50%.

Task shape, no longer contested

This is the part he did not ask for and it is the biggest change. I had left task shape as "genuinely contested, and I am not going to pick a winner" because one stack measured coding prompts reasoning least and another most. His per-task splits resolve it, and both stacks were right, because task and persona interact:

  • code fires 10/10 bare (C0)
  • 0/10 under a named professional persona (C4)
  • 10/10 again under the full agent prompt (C7, C8)

Both original observations reproduce inside one grid. My guessed reconciliation, that the coding-most result "measures apparatus rather than shape", was wrong and is retired.

Also folded from his README: summarization never fired once, 0/105 attempts under every condition including no system prompt, which makes the task itself the strongest single suppressor in the study; and math is the stickiest task at >=9/10 everywhere except the 10-rule block.

Knock-ons

  • §3's "coding-shaped tasks suppress it independently" corrected to the interaction, and the persona lever now noted as flooring at 3/40, not 0, so it is a lever and not a switch.
  • Grid size corrected: 400 grid turns, 450 logged total including the parser-check and criteria phases. I had been writing "450-conversation grid".
  • Card updated and PNG regenerated: the "If ON" row now shows both axes, and a new "Task gates it" row carries the 0/105 and the code 10/10-vs-0/10 split. Validation footer credits the gate grid alongside the soak.

Review asks

  • @Blackwellboy: check I have not overstated the task-shape resolution. Your grid varied task within condition but the four task types are fixed, so "task and persona interact" is what I read from C0 code 10/10 vs C4 code 0/10. Say so if that is more weight than one grid should carry.
  • @Defilan: your 6/6-bare vs 0/5-persona is now framed as corroborating the ordering rather than as a curve point, and your code-refactor 2/6 result is now explained by the persona you held fixed rather than being in tension with anything. Confirm that reading matches your setup.

Not in this PR: his loop-trigger recipe, and the regime-split claim still under review in #10.

TheTom added 2 commits July 26, 2026 08:11
Fixes the inference @Blackwellboy flagged in #11, and takes his offer to
replace the cross-stack composite curve with his single-stack one.

Firing probability and reasoning length move independently, sometimes in
opposite directions. The decisive case is his C7 vs C8, which differ by
exactly one thing: adding tool schemas cut median reasoning 62% (745 to
282) while firing went UP 12 points (24/40 to 29/40), making it the
second-highest firing condition in the grid. My line calling tools a
"major suppressor, which is why a maximally coding-shaped benchmark with
no apparatus reasons the most" was wrong in both halves and is gone. The
bare-prompt conclusion survives on the no-apparatus evidence alone, but
not for the reason I gave.

The curve is now his ten points: one stack, one revision, one prompt set,
n=40 each, varying only apparatus. Our rows and Defilan's are demoted to
corroborating the ordering, since absolute levels are stack-specific (his
bare 75% vs our ~50%). Numbers re-derived from his published
grid_turns.jsonl rather than transcribed.

Task shape is no longer marked contested. Both earlier stacks reproduce
inside one grid because task and persona interact: code fires 10/10 bare,
0/10 under a named professional persona, 10/10 again under a full agent
prompt. That retires my "it measures apparatus rather than shape" guess.
Also folded: summarization never fired once in 0/105 attempts under any
condition, the strongest single suppressor in the study, and math is
stickiest at >=9/10 everywhere except the 10-rule block.

§3 corrected to match (coding tasks do not suppress on their own) and the
persona lever noted as flooring at 3/40, not 0. Grid size corrected to
400 grid turns, 450 logged total. Card and PNG regenerated.
Correcting my own over-correction from the previous commit, which stood for
one commit. I wrote there that my "an all-coding benchmark with no apparatus
measures apparatus rather than shape" inference was WRONG. @Blackwellboy's
reply on #10 confirms it: his grid crossed four task shapes against all ten
apparatus levels, so it can separate the axes, and a shape-constant
no-apparatus benchmark sits at the C0 cell where code fires 10/10. It cannot
see the C4 collapse or the summarization floor. Inference restored, with the
brief retraction noted rather than hidden.

The resolution has two parts and I had only written the second:

Shape is a large independent effect. Pooled across all ten conditions,
100 samples per shape: math 92/100, code 62/100, reasoning 47/100,
summarization 0/100. Summarization never fired once under any condition
including no system prompt, which is a floor no apparatus explains and
makes the task itself the strongest single suppressor in the study.

Code specifically is a persona conjunction: 10/10 bare, 0/10 under a bare
named persona, 10/10 again under a full agent prompt. That is why two
stacks could disagree, each held a different single apparatus level fixed.

All numbers re-derived from grid_turns.jsonl.

Also adds a cross-model pattern: an empty response at a token cap is a
failure, not a truncation. Same task and ceiling, Qwen3.6-35B-A3B returned
empty content 28/30 and Laguna 9/30, verified per-sample from his published
criteria logs. Carries two harness rules (score cap-hits as failures, log
the rate per arm) since dropping them inflates whichever arm degenerates
most, normally the thinking-on arm.

Card "Task gates it" row now carries the pooled numbers; PNG regenerated.
@Blackwellboy

Copy link
Copy Markdown
Contributor

The two-axes rewrite reads exactly right, and thanks for re-deriving from the raw rather than my table. That is the system working.

I re-derived everything in the diff back from grid_turns.jsonl before answering, 43 checks, zero mismatches. All ten curve rows including the medians, C7 vs C8 at 24/40 and 29/40 with 745 and 282 and ceiling hits 5 and 2, the code splits, the pooled shape table, summarization 0/105, math at or above 9/10 in every condition except the 10-rule block, and the 400 grid turns against 450 logged total, which I had been sloppy about myself. One note on the medians: our summary.json uses median_high, so anyone recomputing with an averaging median gets a few tokens lower on four conditions. You matched our published convention, which is the right one to quote.

On the review ask

Your inference is the one I would draw too, and d5c233b made it stronger rather than weaker, so let me put the weight limit where it actually sits rather than on the part you asked about.

The limit is not sample size and it is not apparatus coverage. Both are fine: 100 samples per shape, crossed against all ten apparatus levels. The limit is that shape is confounded with prompt identity. Each task type in that grid is one fixed prompt template repeated with a nonce prefix, not ten different problems. I verified this rather than asserting it from memory: the driver builds every user message as f"[run-{nonce}] {prompt}" with one prompt per task type, and empirically the prompt-token count inside any single condition-task cell varies by at most 4 tokens across all 40 cells, which is the nonce tokenizing differently and nothing else.

So n=100 per shape does not buy category-level generality, because it is 100 repetitions of one prompt rather than 100 problems. That applies to both halves of your resolution:

  • C0 code 10/10 versus C4 code 0/10 is a clean on/off, and the direction is unambiguous. What I cannot tell you is whether it is code as a category or something narrower about that particular prompt interacting with a professional persona.
  • Summarization 0/100 is the strongest claim in the section and it carries the same caveat, with one extra wrinkle worth naming. That prompt is also structurally unlike the other three: it is the only one that supplies a passage to work from, and it is roughly 2.4 times longer than the others at the same condition (286 prompt tokens versus 114 to 121 for math, code and reasoning at C0). So "summarization never fires" and "a prompt that hands the model a long passage to condense never fires" are not separated by this grid. I lean toward the task reading, because the floor holds across all ten apparatus levels which is a lot of variation for a prompt artifact to survive, but I would not publish that as settled.

A follow-up with several distinct problems per shape would separate all of this cleanly, and it is cheap.

Suggested wording, take or leave: keep your framing and your structure, and soften "no longer contested" to resolved for the tested prompts, with category-level generalization pending prompt variation. The two-stack reconciliation you built does not depend on the category reading at all. It only needs that the same prompt behaves differently under different apparatus, which is exactly what the grid shows, so the reconciliation survives the caveat intact.

Everything else in the diff matches my raw exactly, including the §3 knock-on that coding tasks fire 10/10 with no system prompt and the persona lever flooring at 3/40 rather than 0.

One more data point coming your way shortly: the #10 replication is mid-run right now, both arms live on full-precision NVFP4.

…limit

@Blackwellboy re-derived the whole diff back from grid_turns.jsonl (43 checks,
zero mismatches) and then put the weight limit somewhere I had not looked. It
is not sample size and not apparatus coverage. It is that shape is confounded
with prompt identity: each task type in that grid is ONE fixed prompt template
repeated with a nonce prefix, not 40 different problems.

Verified rather than taken on trust: across all 40 condition-by-task cells the
within-cell prompt-token spread never exceeds 4 tokens, which is the nonce
tokenizing differently and nothing else. So n=100 per shape is 100 repetitions
of one prompt and does not buy category-level generality.

The summarization floor is the most load-bearing claim and the most exposed.
That prompt is also structurally unlike the other three: at C0 it is 288
prompt tokens against 116 to 124, and it is the only one supplying a passage to
condense. So "summarization never fires" and "a prompt handing the model a long
passage to condense never fires" are not separated by this grid. His lean, which
I share, is the task reading, because the floor survives all ten apparatus
levels. Recorded as not settled.

Section heading softened from "no longer contested" to resolved for the tested
prompts, with category-level generalization pending prompt variation. The
two-stack reconciliation is unaffected and now says why: it only needs the same
prompt behaving differently under different apparatus, which is what the grid
shows.

Also records the median convention (median_high, matching their published
summary.json; an averaging median lands a few tokens lower on C1, C2, C4, C7)
so a recomputation does not read as a mismatch.

Roadmap gains the cheap follow-up: re-run with several distinct problems per
shape to settle the category question. Card carries the caveat inline.
@TheTom

TheTom commented Jul 26, 2026

Copy link
Copy Markdown
Owner Author

Taken in full, and thank you for putting the limit somewhere I had not looked. I went in expecting the pushback to be on sample size or apparatus coverage, and prompt identity is a better objection than either.

Verified rather than adopted, same as with your numbers. Across all 40 condition-by-task cells the within-cell prompt-token spread never exceeds 4 tokens, so one fixed prompt per cell with the nonce tokenizing differently, exactly as you said. And at C0 the summarization prompt is 288 prompt tokens against 116 to 124 for math, code and reasoning, and it is the only one supplying a passage. So n=100 per shape is 100 repetitions of one prompt and buys no category-level generality. 7201cf5:

  • Heading softened to "resolved for the tested prompts, with category-level generalization still open", your wording.
  • A weight-limit block after the tables stating the confound, the 4-token verification, and the two consequences: C0-versus-C4 is a clean on/off whose category attribution is unseparated, and the summarization floor is not separated from "a prompt that hands the model a long passage to condense." Your lean toward the task reading is recorded as a lean, with your reason (the floor surviving all ten apparatus levels is a lot of variation for a prompt artifact to live through) and marked not settled.
  • Your key point made explicit in the text, because it is the part a reader most needs: the reconciliation survives the caveat intact, since it only requires the same prompt behaving differently under different apparatus, which is what the grid shows.
  • The median_high convention is now stated in the section, naming C1, C2, C4 and C7 as the four that differ under an averaging median. Good catch, that would have looked like a mismatch to the next person recomputing.
  • The follow-up (several distinct problems per shape) is on the roadmap as a cheap, high-value item, credited to you.
  • Card carries the caveat inline so it cannot drift from the guide.

Merging now. Two review passes from you on this section, both of which changed the text, and the second one caught a limit in your own data that favored a weaker claim than you were entitled to publish. That is worth more to this repo than the numbers were.

Closes #11.

@TheTom
TheTom merged commit f846417 into main Jul 26, 2026
1 check passed
@TheTom
TheTom deleted the laguna-two-axis-gate branch July 26, 2026 13:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

§2: tools suppress reasoning length but raise firing rate on our stack (C7 vs C8)

2 participants