Laguna §2: the thinking gate is two axes, not one dose curve (closes #11)#12
Conversation
Fixes the inference @Blackwellboy flagged in #11, and takes his offer to replace the cross-stack composite curve with his single-stack one. Firing probability and reasoning length move independently, sometimes in opposite directions. The decisive case is his C7 vs C8, which differ by exactly one thing: adding tool schemas cut median reasoning 62% (745 to 282) while firing went UP 12 points (24/40 to 29/40), making it the second-highest firing condition in the grid. My line calling tools a "major suppressor, which is why a maximally coding-shaped benchmark with no apparatus reasons the most" was wrong in both halves and is gone. The bare-prompt conclusion survives on the no-apparatus evidence alone, but not for the reason I gave. The curve is now his ten points: one stack, one revision, one prompt set, n=40 each, varying only apparatus. Our rows and Defilan's are demoted to corroborating the ordering, since absolute levels are stack-specific (his bare 75% vs our ~50%). Numbers re-derived from his published grid_turns.jsonl rather than transcribed. Task shape is no longer marked contested. Both earlier stacks reproduce inside one grid because task and persona interact: code fires 10/10 bare, 0/10 under a named professional persona, 10/10 again under a full agent prompt. That retires my "it measures apparatus rather than shape" guess. Also folded: summarization never fired once in 0/105 attempts under any condition, the strongest single suppressor in the study, and math is stickiest at >=9/10 everywhere except the 10-rule block. §3 corrected to match (coding tasks do not suppress on their own) and the persona lever noted as flooring at 3/40, not 0. Grid size corrected to 400 grid turns, 450 logged total. Card and PNG regenerated.
Correcting my own over-correction from the previous commit, which stood for one commit. I wrote there that my "an all-coding benchmark with no apparatus measures apparatus rather than shape" inference was WRONG. @Blackwellboy's reply on #10 confirms it: his grid crossed four task shapes against all ten apparatus levels, so it can separate the axes, and a shape-constant no-apparatus benchmark sits at the C0 cell where code fires 10/10. It cannot see the C4 collapse or the summarization floor. Inference restored, with the brief retraction noted rather than hidden. The resolution has two parts and I had only written the second: Shape is a large independent effect. Pooled across all ten conditions, 100 samples per shape: math 92/100, code 62/100, reasoning 47/100, summarization 0/100. Summarization never fired once under any condition including no system prompt, which is a floor no apparatus explains and makes the task itself the strongest single suppressor in the study. Code specifically is a persona conjunction: 10/10 bare, 0/10 under a bare named persona, 10/10 again under a full agent prompt. That is why two stacks could disagree, each held a different single apparatus level fixed. All numbers re-derived from grid_turns.jsonl. Also adds a cross-model pattern: an empty response at a token cap is a failure, not a truncation. Same task and ceiling, Qwen3.6-35B-A3B returned empty content 28/30 and Laguna 9/30, verified per-sample from his published criteria logs. Carries two harness rules (score cap-hits as failures, log the rate per arm) since dropping them inflates whichever arm degenerates most, normally the thinking-on arm. Card "Task gates it" row now carries the pooled numbers; PNG regenerated.
|
The two-axes rewrite reads exactly right, and thanks for re-deriving from the raw rather than my table. That is the system working. I re-derived everything in the diff back from On the review askYour inference is the one I would draw too, and The limit is not sample size and it is not apparatus coverage. Both are fine: 100 samples per shape, crossed against all ten apparatus levels. The limit is that shape is confounded with prompt identity. Each task type in that grid is one fixed prompt template repeated with a nonce prefix, not ten different problems. I verified this rather than asserting it from memory: the driver builds every user message as So n=100 per shape does not buy category-level generality, because it is 100 repetitions of one prompt rather than 100 problems. That applies to both halves of your resolution:
A follow-up with several distinct problems per shape would separate all of this cleanly, and it is cheap. Suggested wording, take or leave: keep your framing and your structure, and soften "no longer contested" to resolved for the tested prompts, with category-level generalization pending prompt variation. The two-stack reconciliation you built does not depend on the category reading at all. It only needs that the same prompt behaves differently under different apparatus, which is exactly what the grid shows, so the reconciliation survives the caveat intact. Everything else in the diff matches my raw exactly, including the §3 knock-on that coding tasks fire 10/10 with no system prompt and the persona lever flooring at 3/40 rather than 0. One more data point coming your way shortly: the #10 replication is mid-run right now, both arms live on full-precision NVFP4. |
…limit @Blackwellboy re-derived the whole diff back from grid_turns.jsonl (43 checks, zero mismatches) and then put the weight limit somewhere I had not looked. It is not sample size and not apparatus coverage. It is that shape is confounded with prompt identity: each task type in that grid is ONE fixed prompt template repeated with a nonce prefix, not 40 different problems. Verified rather than taken on trust: across all 40 condition-by-task cells the within-cell prompt-token spread never exceeds 4 tokens, which is the nonce tokenizing differently and nothing else. So n=100 per shape is 100 repetitions of one prompt and does not buy category-level generality. The summarization floor is the most load-bearing claim and the most exposed. That prompt is also structurally unlike the other three: at C0 it is 288 prompt tokens against 116 to 124, and it is the only one supplying a passage to condense. So "summarization never fires" and "a prompt handing the model a long passage to condense never fires" are not separated by this grid. His lean, which I share, is the task reading, because the floor survives all ten apparatus levels. Recorded as not settled. Section heading softened from "no longer contested" to resolved for the tested prompts, with category-level generalization pending prompt variation. The two-stack reconciliation is unaffected and now says why: it only needs the same prompt behaving differently under different apparatus, which is what the grid shows. Also records the median convention (median_high, matching their published summary.json; an averaging median lands a few tokens lower on C1, C2, C4, C7) so a recomputation does not read as a mismatch. Roadmap gains the cheap follow-up: re-run with several distinct problems per shape to settle the category question. Card carries the caveat inline.
|
Taken in full, and thank you for putting the limit somewhere I had not looked. I went in expecting the pushback to be on sample size or apparatus coverage, and prompt identity is a better objection than either. Verified rather than adopted, same as with your numbers. Across all 40 condition-by-task cells the within-cell prompt-token spread never exceeds 4 tokens, so one fixed prompt per cell with the nonce tokenizing differently, exactly as you said. And at C0 the summarization prompt is 288 prompt tokens against 116 to 124 for math, code and reasoning, and it is the only one supplying a passage. So n=100 per shape is 100 repetitions of one prompt and buys no category-level generality.
Merging now. Two review passes from you on this section, both of which changed the text, and the second one caught a limit in your own data that favored a weaker claim than you were entitled to publish. That is worth more to this repo than the numbers were. Closes #11. |
Closes #11. @Blackwellboy caught a bad inference in §2 and offered a better version of the curve. Both taken.
The correction
The sentence he quoted had already been removed in 221a311 during the dose-curve retraction, but the substance was still live elsewhere in §2, because I was still presenting the compression numbers (3536 to 745 to 282) as a single suppression progression without saying that firing went up across the last step. So this is not a no-op.
His C7 vs C8 differ by exactly one thing, tool schemas in the request:
Reasoning length down 62%, firing up 12 points, second-highest firing condition in his whole grid. I re-derived every number in this PR from his published
grid_turns.jsonlrather than transcribing his table, and all ten conditions match exactly.§2 now states the two axes separately, with C7/C8 as the worked example of them disagreeing. His point about why it matters is the part I have quoted into the guide in substance: collapsing firing and length into one "suppressor" axis is what made the persona-times-task gate confusing for three people independently.
The curve
Taking his offer. The three-stack composite is gone, replaced by his ten single-stack points (one rev, one prompt set, n=40 each, varying only apparatus). He was right that a curve assembled from three stacks with three different prompt sets is the weakest form of the argument. Our rows and @Defilan's are demoted to corroborating the ordering, with a note that absolute levels are stack-specific, since his bare rate is 75% against our ~50%.
Task shape, no longer contested
This is the part he did not ask for and it is the biggest change. I had left task shape as "genuinely contested, and I am not going to pick a winner" because one stack measured coding prompts reasoning least and another most. His per-task splits resolve it, and both stacks were right, because task and persona interact:
Both original observations reproduce inside one grid. My guessed reconciliation, that the coding-most result "measures apparatus rather than shape", was wrong and is retired.
Also folded from his README: summarization never fired once, 0/105 attempts under every condition including no system prompt, which makes the task itself the strongest single suppressor in the study; and math is the stickiest task at >=9/10 everywhere except the 10-rule block.
Knock-ons
Review asks
Not in this PR: his loop-trigger recipe, and the regime-split claim still under review in #10.