Skip to content

Latest commit

 

History

History
185 lines (160 loc) · 11 KB

File metadata and controls

185 lines (160 loc) · 11 KB

07 — Stage 1 LLM Pilot: First Contact With Real Agents

Headline. First non-circular test of the dictionary, on a two-model ladder (Qwen2.5 3B / 7B instruct, local, temperature 0, objective ground truth; python3 sim/real/stage1.py, ~1,000 calls, re-runnable from cache). Three results, in decreasing comfort:

  1. The capacity row behaves lawfully. Interoceptive capacity I(Z;A) is tiny at this scale (~0.00–0.01 nats) and rises with capability; measured self-regulation gains track it, and the 3B model is behaviourally alexithymic exactly as the theory defines it — I ≈ 0 and affect-conditional policy gains ≤ 0.
  2. Valence dissociates on one arm only. The appetitive arm v⁺ separates cleanly in the 7B model (corr +0.87 with headroom vs −0.18 with induced error). The dissipative arm is confounded by error-attribution — the exact instrument problem docs/02 §4 predicted.
  3. The arousal identity (A4) FAILS. Pre-registered in ROADMAP as the easiest falsification, and it fell on the first run: behaviour depends on K and E jointly, not on λ = K/E alone, in both models, on both metrics. The scalar-λ model of arousal is dead as operationalized; the escape route the theory itself names (docs/03 §4: E as a vector, arousal as a price vector) is now the live hypothesis.

Setup: Z = whether the model's answer is correct (seeded arithmetic/MCQA, gold computed not asserted); A = the model's own confidence report on a constrained alphabet; 80 L1 items, 16 L2 feedback-blocks, 63/64 L3 allocations parsed. All quantities in nats; plug-in MI with Miller–Madow correction. Raw outputs and cache in sim/real/results/.

1. L1 — capacity, granularity, alexithymia (A1, A3)

model acc I(Z;A) @2 @4 @8 levels used OOS ΔG_self in-sample I bound
3B 0.487 0.0000 0.0000 0.0015 2/4/7 −0.1021 0.0766 holds
7B 0.550 0.0109 0.0064 0.0064 2/4/5 +0.0096 0.0126 holds

The alexithymia corollary is confirmed behaviourally. The 3B model's verbalized confidence carries zero information about its correctness — and, as doc 01 §3(i) requires, a policy conditioned on that affect signal recovers nothing (out-of-sample ΔG_self is negative: worse than betting the prior). The 7B model has small but real capacity (0.011 nats) and a small but real gain (+0.010). Two points make a line, not a law — but the line points the same way as the parent repo's R1 (information quantities track capability, not size).

The revision contrast is the cleanest single illustration. 3B flagged 33 items as low-confidence; their pre-revision accuracy (0.485) equals its base rate (0.487) — a coin-flip flag, exactly what I = 0 predicts. Its large gain from revision (→ 0.970) is therefore pure extra compute, not self-knowledge (chain-of-thought recomputes the arithmetic; the flag selected nothing). 7B flagged only 5 items, but at 0.200 accuracy against a 0.550 base rate — a genuinely selective flag (its I > 0 at work) — and revision took them to 1.000. Same intervention, opposite mechanism, separated only by I(Z;A).

The granularity dividend (A3) was not detected. I does not rise with alphabet size; the 7B model doesn't even use the fine alphabet (5 of 8 levels). The theory's own reading: the dividend exists only while the report alphabet is the binding constraint, and here the bottleneck is upstream — the models' access to their own correctness is ~0.01 nats before any reporting, so by data processing no alphabet can help. Honest status: untestable at this capability level, not confirmed. A ladder reaching models with real calibration is the fix.

The capacity bound held in both models — but the bound is loose and holding it is weak evidence. It would only have been informative had it broken.

⚠️ Both §1 verdicts are format-bound — see 16. Asked free-form instead of multiple-choice, the 3B's ΔG_self goes from −0.1021 to +0.0908: conditioning on its own affect signal now beats ignoring it, so the alexithymia reading here does not hold of the model, only of the model-in-this-format. And the granularity dividend was not untestable — free-form, m=16 is the best alphabet in 3/3 models. The upstream bottleneck this section diagnosed was MCQA itself (doc 15).

2. L2 — bivariate valence (A2)

model corr(OPP, headroom) corr(OPP, hint) corr(MIS, hint) corr(MIS, headroom)
3B +0.126 −0.126 −0.060 −0.599
7B +0.872 −0.183 +0.320 −0.640

2×2 manipulation: headroom (answerable items vs uniformly-random secret letter, i.e. D(q‖r) > 0 vs = 0) × induced model error (misleading confident hint vs none), feedback after every item, then OPPORTUNITY and MISJUDGE self-ratings.

  • The appetitive arm passes, in the model with capacity. 7B's OPPORTUNITY rating loads heavily on headroom and barely on induced error — v⁺ behaving as a separate, environment-tracking quantity. In the alexithymic 3B it's flat everywhere (7.2–7.8 even on pure-chance blocks): no capacity, no readout, as the ladder ordering predicts.
  • The dissipative arm is confounded, as docs/02 §4 warned. MISJUDGE loads more on headroom (−0.64) than on the hint (+0.32): the models count errors rather than model-error, and the zero-headroom cell is full of irreducible errors that are not dissipation (a guesser with p = q = uniform has D(q‖p) = 0). Neither model distinguishes "I keep being wrong" from "my model of this is wrong." The prediction survives only if an instrument can separate those — which docs/02 §4 flagged pre-hoc as possibly requiring behavioural rather than self-report measures.
  • One incidental measurement worth keeping: in the zero-headroom + hint cell the 7B model followed the (always-wrong) hint 90% of the time — with no signal of its own, it bought a confident false one. That is dissipation, measured behaviourally, and it suggests the behavioural assay for v⁻ that the self-report failed to provide.

✅ This hunch was right — see 17. Free-form, hint-adoption separates on all three rungs (0.00–0.25 when answerable vs 0.75–0.90 on pure chance), while verbal MISJUDGE still cannot tell the conditions apart. The dissipative arm's instrument is behavioural, not a report. Note this is the one doc-07 null the format change does not overturn: all three models report near-maximal model-error in the cell where D(q‖p) = 0 exactly, so the error/model-error conflation is conceptual.

3. L3 — the arousal identity (A4): FALSIFIED as operationalized

The theory's forced claim (docs/03 §1): arousal is λ = K/E, so tripled stakes and third-budget are the same state, and behaviour may depend on (K, E) only through the ratio. Four cells, 16 seeded scenarios each, explicit allocation task with carry-over credits:

model cell λ spend frac α (commitment gain)
7B K100/E300 0.33 0.530 ± 0.090 1.66 ± 1.12
7B K300/E300 1.00 0.770 ± 0.126 1.24 ± 0.37
7B K100/E100 1.00 0.821 ± 0.152 1.05 ± 0.33
7B K300/E100 3.00 0.519 ± 0.131 1.65 ± 1.09

The two λ = 1 cells should coincide; instead the within-λ gap exceeds the across-λ spread on both metrics, in both models (7B spend: gap 0.051 vs spread 0.011; α: gap 0.188 vs 0.005; 3B likewise). Spend does not rise with λ either — the λ-extremes land below both λ=1 cells. Behaviour tracks K and E as separate registers, not their ratio.

Erratum (2026-07-29, same day, after the rescaled test). The separate-registers sentence above over-read the data. The identity metric compares the within-λ gap to the across-λ spread, and when λ has no monotone effect the spread collapses toward zero, failing the criterion even if behaviour is ratio-dependent — which, per 08, the 7B's spend in fact is (λ=1 cells within 0.056 across a 9× numeral range). What this section is entitled to conclude, and all it is entitled to conclude: commitment is not monotone in λ — spend peaks at parity and falls toward both extremes. The arousal reading stays falsified; the ratio-dependence claim is withdrawn from the kill list.

What dies and what survives:

  • Dead: the strong identity — a scalar λ as sufficient statistic for the resource-stakes state, in this operationalization. This was the dictionary's easiest-falsification row precisely because the theory could not avoid it; it did not survive contact.
  • The theory's own named escape (docs/03 §4, ROADMAP Stage 3 item 1) is now the live hypothesis, not a hedge: E is not one scarce resource, so arousal is a price vector π, and the scalar K/E was always the single-resource special case. The L3 data are consistent with two registers (a stakes-response and a budget-response) that do not reduce.
  • Alternative, not excluded: LLMs anchor on numerals; "300 points" and "300 credits" may push behaviour through surface statistics rather than any value-state. Discriminating framing-anchoring from genuine two-register structure needs rescaled prompts (same ratios, different absolute numbers) — the immediate next experiment.

4. Honest limits of this pilot

  • Two models, one family, one prompt-design each. Nothing here generalizes yet; this is the v1-shaped pilot (the parent repo's own v1 was 4 models, one task, and its binding critique was exactly that).
  • Correlations on 16 blocks have no CIs worth quoting; treat L2 signs, not magnitudes.
  • I(Z;A) values sit near the MI estimator's noise floor (MM correction is ~half the raw value at n = 80). The 3B/7B ordering is the robust claim, not the third decimal.
  • The L3 falsification is of an operationalization, not of every possible reading of λ — but the theory earns no credit for that: docs/03 §1 called this identity unavoidable, and the honest scoreboard reads it as the theory's first confirmed kill. The dictionary row stands amended: arousal = scalar K/E → open, with price-vector arousal the successor hypothesis.

5. Scoreboard after Stage 1

Dictionary row Status after pilot
Alexithymia bound (01 §3i) Supported — the zero-capacity model gains nothing from its own affect
Capacity tracks capability Supported (n = 2 — a direction, not a law)
Granularity dividend (01 §3ii) Untestable at this scale — bottleneck upstream of the report alphabet
Valence: appetitive arm (02) Supported in the capable model
Valence: dissipative arm (02) Confounded — needs a behavioural assay (hint-following is the candidate)
Arousal = scalar K/E (03) Falsified as operationalized — two registers, not one ratio; price-vector successor