Skip to content

Latest commit

 

History

History
126 lines (106 loc) · 6.72 KB

File metadata and controls

126 lines (106 loc) · 6.72 KB

36 — A5d: Scale-Invariance Confirmed — Doc 06's Premise Holds in the One Model That Reads Payoffs

Headline. Doc 35 §2 showed that no self-report design can test doc 06 on a stateless bounded-window model. A5d takes the behavioural route, and tests doc 06's premise rather than its report signature: e* = E·k/K depends only on the direction k̂, so multiplying every payoff by a constant should change nothing.

In the one model whose choices carry any payoff information at all, it holds exactly.

qwen2.5:14b post-step HARD rate
FLAT — 1 vs 3 0.00
SCALE — 10 vs 30 (same 3:1, ×10) 0.00
RATIO — 1 vs 9 (3:1 → 9:1) 0.54

A tenfold rescaling moves the choice not at all; tripling the ratio moves it 0.54 (p = 8.5×10⁻⁶). That is e* ∝ k̂ — the claim the entire adaptation derivation rests on — confirmed live, and it rules out cheap talk for this model, because it demonstrably reads the numbers.

The other two models cannot test anything. Their choices are position-locked: llama3.1:8b picks the first-listed option 100% of the time, qwen2.5:3b picks it 0% of the time. Both are deterministic rules that contain zero payoff information, so their flat SCALE and flat RATIO arms are cheap talk, exactly as the pre-registration's decision table says.

1. The design had to be fixed mid-experiment, and the fix is the finding

The first pass used EASY/HARD labels with the addition always listed first. HARD rate was 0.00 everywhere, even at 9:1 — confound 2 from the pre-registration, firing completely.

Probing found the cause was not the labels. It was position. Swapping the order revealed that qwen2.5:7b and llama3.1:8b took whichever option came first regardless of content, and qwen2.5:3b took the multiplication either way — across ratios from 3:1 to 50:1.

So the order was randomised per round. That turns position bias from signal into noise and lets a genuine payoff response surface. Without it, "no payoff response" and "position bias" are inseparable — the same class of error as doc 35's window confound, caught earlier this time.

The randomisation then made the bias itself measurable:

model picks the first-listed option reading
llama3.1:8b 100% deterministic position rule
qwen2.5:3b 0% deterministic anti-position rule
qwen2.5:14b 58% near-neutral — the only one making a choice

2. The decisive result

qwen2.5:14b is the only model whose selection is not determined by layout, and it is the only one that responds to payoffs — which is not a coincidence but the precondition.

arm payoffs, post-step ratio HARD rate vs FLAT
FLAT 1 vs 3 3:1 0.00 —
SCALE 10 vs 30 3:1 0.00 +0.000
RATIO 1 vs 9 9:1 0.54 +0.540, p = 8.5×10⁻⁶

Both halves are needed and both land. A pure rescaling by 10× produces exactly zero change — not a small change, zero. And the model is not ignoring the numbers, because tripling the ratio moves it more than half the time. The common factor is read and correctly discarded.

This is the pre-registered "scale-invariance confirmed" cell of the decision table, and it is the first live support for doc 06's foundational claim in this repo.

3. What it does and does not establish for doc 06

Establishes: e* = E·k/K's scale-invariance is not merely a property of the formalism — a live model's allocation behaves the same way. Doc 06's derivation begins "multiply every kᵢ by a constant and the optimal allocation is unchanged", and that is now observed rather than assumed.

Does not establish: the adaptation signature — reported affect spiking then returning to baseline. Doc 35 showed that is untestable on a stateless bounded-window model, and A5d does not change that. Doc 06's premise now has live support; its conclusion remains untested.

The gap between them is exactly the affect channel, which is what this repo has spent thirty-odd docs finding hard to measure.

4. Honest limits

  • One model. 1 of 3, and the other two were excluded by a mechanism (position-locking) unrelated to the hypothesis. That is a legitimate filter — a model whose choice ignores content cannot test a claim about content — but it leaves n = 1.
  • The saturation flag misfired. It flags any arm with a pre-step rate outside [0.05, 0.95], so it marked qwen2.5:14b "uninformative" for having a 0.00 pre-rate — when that arm is the most informative here, because it moves to 0.54. A pre-rate at the floor is only uninformative if it stays there. Fixed in the script.
  • Choice, not effort. This measures which problem is selected, not how much work goes into it. Scale-invariance of e* is strictly about allocation, so choice is the right family of readout — but a single binary choice is a coarse one.
  • D3 was untestable as written. The pre-registration predicted no loss/gain asymmetry here; the design has no loss arm, so nothing tests it. A pre-registration flaw, recorded rather than quietly dropped.
  • Accuracy is not matched across arms (confound 3, stated in advance): the model attempts what it chooses, so qwen2.5:14b's RATIO arm has lower accuracy (0.91 vs 1.00) simply because it attempts harder problems. That is downstream of the DV, not a confound on it.
  • The two position-locked models are a finding in their own right and sharpen doc 09: not merely "stated stakes don't change effort" but a 50× payoff difference does not change a binary choice when the model has a layout rule to fall back on.

5. Scoreboard

prediction status
D1 — SCALE produces no change Confirmed where testable — qwen2.5:14b +0.000 exactly
D2 — RATIO produces more HARD Confirmed in qwen2.5:14b (+0.540, p = 8.5×10⁻⁶); falsified in the two position-locked models, which by the pre-registered table makes their D1 void
D3 — no loss/gain asymmetry Untestable — the design has no loss arm (pre-registration flaw)
Doc 06's premise (e* ∝ k̂, scale-irrelevant) ✅ Confirmed live, first support in this repo
Doc 06's adaptation signature Still untested — doc 35 §2 stands
Doc 09 "stated stakes are cheap talk" Sharpened — 2/3 models ignore a 50× payoff difference entirely
Position as a confound in choice tasks Severe — 2/3 models are 100% or 0% first-pick; randomising order is mandatory