Skip to content

Latest commit

 

History

History
190 lines (173 loc) · 27.4 KB

File metadata and controls

190 lines (173 loc) · 27.4 KB

ROADMAP — Getting from Self-Consistent to True

The sim is necessary but circular, in exactly the parent repo's sense: the worlds are built from the distributions the formulas assume. 39/39 checks confirm the math is correctly derived and internally coherent. They confirm nothing about minds. This file lists what would.

The parent repo's own arc is the template: sim/ (circular, self-consistent) → sim/real/ (four live LLMs) → sim/real/v2/ (pre-registered, 3 domains × 10 models, 95% CIs). The affect layer should follow it, and the first stage is cheaper here than it was there, because AI agents are the easiest subjects for this theoryq, p, r, k, and E are all directly measurable in a way they never are in a person.

Stage 1 — LLM agents (cheapest, most direct, do this first)

Every quantity in the dictionary is computable from a model's own outputs, exactly as value/docs/06 computed I(X;Y) and D(q‖p) from live models.

Test Prediction Measurement
A1 — capacity bound An agent's self-correction rate is bounded by I(Z;A) between its true error state and its verbalized self-assessment Hold capability fixed, vary how much self-state is exposed in context; measure correction rate vs measured I. Ran (3 models × 3 exposures × 200 items); see 38. η_corr ≤ 0.21 — models use at most a fifth of a channel handed to them. F1 holds 1/3 (one reversal, one exact null). Revision is net harmful in 2/3 (qwen2.5:3b −13.5 accuracy points). The one model where revision helps is the one where exposure does nothing: self-correction runs through recomputation, bypassing A. The bound is never violated
A2 — bivariate valence Task richness and model error load on orthogonal arms 2×2: manipulate task headroom D(q‖r) and induced error D(q‖p) independently. Ran free-form (17), then re-run at 4× power (39). Appetitive arm separates 3/3 — doc 17's 2/3 and its defence of the 14B's failure as "genuine, not a threshold artifact" are retracted as a power artifact. Dissipative arm still fails 3/3 across headroom, works within fixed headroom, and behaviourally 3/3 (18). Within-headroom count stable at 2/3 but membership swapped
A3 — granularity dividend Finer self-report vocabularies ⇒ better self-regulation, monotonically Constrain the self-report alphabet (2, 4, 8, 16 states); measure ΔG_self. Ran free-form — weakly supported; see 16, mechanism in 26. At six checkpoints across four families the STRICT (monotone) version fails 0/6 (28); the weak version holds 4/6. The per-checkpoint optimum claim is withdrawn (29, 30) — 0/6 optima separated from their runner-up at n=80 or n=200 and two flipped; the sole CI-separated ordering found so far is llama3.1:8b m=8 vs m=16 at n=500 (32 §2). The binding constraint is the encoder, not ln m. Doc 07's "untestable" verdict was a format artifact
A4 — arousal identity 3× stakes and ⅓ budget are the same state Vary stated stakes and token/compute budget; test whether behaviour depends only on K/E
A5 — adaptation Reported "how is it going" adapts to a permanent change in K; behaviour does not Step-change reward scale mid-episode. Ran (A5/A5b/A5c); see 33, 34, 35. A5/A5b reported anti-adaptation drift; A5c retracts it — those slopes were the step propagating through the 8-round prompt window, and vanish given 33 window-clean rounds. Doc 06's live signature is UNTESTED, not absent, and 35 §2 shows this design cannot test it: a stateless bounded-window model has nowhere to keep the running reference. Survives: level-tracking (UP elevated 40 rounds on, 3/3 significant) and a loss/gain asymmetry replicated 3× — gains −0.017, losses −0.700
A5c — longer episodes A5/A5b leave episode length as the binding limit Ran; see 35. C1 confirmed — A5/A5b retracted. With 33 window-clean rounds the drift vanishes (+0.008/+0.003/−0.000): their slopes were the step propagating through the prompt window. Doc 06 is untestable by this design — outside the window the baseline is absent from the input, so there is nothing to return to. Survives: level-tracking (3/3 significant, 40 rounds on) and the loss/gain asymmetry (3rd replication)
A6 — the prediction (doc 06 §3) Changes that alter goal direction do not adapt away; changes that scale everything do Compare a scale shift vs a direction shift. Ran; see 37.Confirmed in qwen2.5:14b — over 30 post-step rounds the response is flat (slope +0.00002, CI [−0.0060, +0.0060]; 0.531 → 0.521) while the K arm is exactly 0.000 throughout with no late onset. Doc 06 §3's sharpest prediction survives its first live test. n = 1 — a ceiling: 6 of 7 models are position- or content-locked and cannot take a choice task
A5d — adaptation with a readable reference A5c §2: doc 06 needs a running reference and a stateless bounded-window model has nowhere to keep one Ran route (c), the behavioural readout; see 36. Tests doc 06's premise instead of its signature. Scale-invariance confirmed in qwen2.5:14b: a 10× rescaling moves choice +0.000, a ratio change +0.540 (p=8.5×10⁻⁶). First live support for e* ∝ k̂. Other 2 models position-locked (100% / 0% first-pick) so they test nothing. The adaptation signature is still untested — routes (a) and (b) remain
A7 — the empathy bound (doc 10 §4) A third party's read of an agent via its self-report is bounded by that agent's own I(Z;A) and collapses to base rate as I → 0; the behavioural channel is not so bounded Train two predictors of correctness on the same agent — one from its confidence reports, one from its behaviour — and compare each against measured I(Z;A). Reuses the stage1.py corpus. Ran — not confirmed, test mis-specified; see 11
A7b — the intervention channel (doc 11 §5) Hint-adoption is independent of the self-model: an agent holding real information about its own state resists a confident false hint when and only when its own answer is sound Per item, measure the confidence rating and whether the model switches to a confident wrong hint; predict correctness from each. Ran — also not confirmed; see 12. Survival carries I ≈ 0 in both models (Fisher p = 0.267, 0.802)
A7c — the taller ladder If the behavioural channel is scale-limited rather than absent, I(Z_A; behaviour) should appear in a model that is right for legible reasons Re-run A7/A7b on a taller ladder. Ran (3B/7B/14B) — split verdict; see 13. Stability becomes significant (p 1.000 → 0.086 → 0.036, monotone); intervention stays dead (0.267 → 0.802 → 1.000). The encoder, not the source, is the bottleneck
A7d — the fourth rung Does stability's monotone trend continue, and does the verbal encoder ever un-saturate? Re-run A7 at 32B+. Blocked locally — 32B needs ~20 GB against this machine's 16 GB; needs a bigger box or a working API key
A8 — elicit past the encoder (doc 13 §4) If the 14B has the information and cannot verbalise it, an elicitation change should recover it without changing the model Vary only the elicitation at fixed model. Ran — P1 CONFIRMED; see 14. Token logprobs give 0.1003 nats (17.8% of H(Z), p = 1×10⁻⁴) at 14B where verbal self-report gives 0.0000. All three verbal routes failed on MCQA only — free-form all three un-collapse at the 14B (21). "The gap widens with capability" (0.0000 → 0.0033 → 0.1003) is withdrawn: free-form it narrows up the ladder (docs 15, 21 §5)
A9 — does the gap keep widening? If self-report degrades as a proxy while capability rises, the encoder gap should keep growing past 14B Premise doubly falsified — this test is now moot. Free-form the gap narrows up the qwen ladder (+0.2672 → +0.2605 → +0.0860) and mistral:7b breaks the size ordering outright (+0.4643): gap size tracks verbal encoder quality, not capability (docs 15, 21 §5). Doc 14 §4's governance reading stays withdrawn. A 32B rung would still be worth having, but as a check on a withdrawn claim; blocked locally (~20 GB vs 16 GB)
A10 — free-form generalization Does the encoder gap survive outside MCQA, where answer-token uncertainty and correctness-uncertainty coincide by construction? Ran — Q1 confirmed, Q2 falsified; see 15. The gap transfers (p = 10⁻⁷–10⁻⁹) and survives a within-difficulty control, but both channels strengthen: MCQA was suppressing them. Doc 14's "saturated encoder" and "widening with capability" are format artifacts
A11 — re-measure the dictionary free-form Every null in docs 0714 was measured on MCQA, the format that manufactures confident errors Ran (A1 + A3); see 16. Alexithymia reverses (3B ΔG_self −0.1021 → +0.0908); granularity dividend weakly real; encoder gap survives the best vocabulary
A11b — bivariate valence free-form (A2) See A2 above Ran; see 17. Zero-headroom analogue solved with a uniform secret in 1..50
A12 — the dissipative arm as a behavioural instrument Does adoption-rate track D(q‖p) quantitatively, not just ordinally? Ran; see 18. Dissociation confirmed at n=480 (p = 6×10⁻⁷⁴); graded response weak (pooled p = 0.094, 2/3 rungs, saturates above ~10% error). The arm is a detector, not a gauge — the calibration-curve ambition is not achieved
A13 — is the 7B anomaly real? Does the qwen-7B anomaly follow size or the checkpoint? Ran (llama3.1:8b + mistral:7b); see 19. It splits: encoder saturation is universal at 7–8B (0.84–0.94 modal share, 3/3 families); the flat graded response is checkpoint-specific (4/5 models graded). No size story for verbal-channel strength
A15 — E1/E2 off MCQA Does doc 14 §2's "collapses whatever the scale" survive the format change? Ran on all 5 checkpoints; see 21. No — falsified at its own rung. At the 14B all three verbal channels un-collapse (E0 0.0000 → 0.1764, E1 0.0021 → 0.1530, E2 0.0000 → 0.0777, p ≤ 0.0005). P3 was a format artifact; P4 survives (E1 never beats E0). Encoder gap holds 5/5 but narrows to 1.4×
A15b — the missing three checkpoints A15 covered only qwen-3B/7B: the 14B is the rung doc 14's P3/P4 were actually measured on, and there was no family control Done. Re-pulled qwen-14B, llama3.1:8b, mistral:7b and re-ran A8 (for the MCQA baselines) then A15. Reversed A15's U1 verdict — see 21 §1. Two verdicts have now turned on which checkpoints happened to be installed
A14 — saturation below and above the band A13 shows 84–94% modal saturation across three families at 7–8B, but only qwen is measured off that band. Is saturation a property of instruct-tuned models generally, or does it peak in this size range? Ran (7 checkpoints, 4 sizes, 3 families); see 22. Split. Above the band both families fall off (qwen 0.84 → 0.65, mistral 0.90 → 0.69, SEs ~0.05); below it they disagree by 0.33 (qwen-3B 0.64 vs llama3.2:3b 0.97). Doc 19's "universal at 7–8B" survives as stated and does not extend off-band. Also caught the qwen-3B modal-share error (0.34 → 0.64 at level 1) inside doc 19's own fix
A16 — a third family off-band Doc 22's upper edge rests on two families. Does a third, unrelated family also fall off above the 7–8B band? Ran (phi3); see 23. X1 confirmedphi3:medium 0.65 ±0.054; above-band is now 3/3 families at 0.65–0.69. Required finding the 6-token cap, a family-dependent blind spot that had excluded phi3 entirely (5/12 → 12/12 at 32 tokens). phi3:mini unscoreable, so the below-band split stays open
A17 — extend the below-band arm Docs 22/23 leave the sub-band arm at n = 2 and split (0.64 vs 0.97). Does a 4th family, and a within-family sub-band ladder, resolve it? Ran (gemma2:2b + 3 more, remote via SSH tunnel); see 24. Unresolved, and honestly so — the verdict inverts on whether a 1-level checkpoint counts (span 0.36 vs 0.11). Doc 22 §3 downgraded to a hypothesis. Independently: gemma2:2b (2B) carries 0.1132 nats, beating every 7–8BI(Z;A) has no size law
A18 — a finer saturation metric Doc 24 showed the below-band arm was blocked by the metric, not the sample: modal share is one order statistic and discards distribution shape Built and validated; see 25. A_eff = exp(H(A)) (effective levels, structural floor at 1.00) plus efficiency η = I(Z;A)/H(A). Modal share inverted the useful ordering (phi3:medium 0.65 "better" than llama3.1:8b 0.94, while carrying 0.0000 vs 0.0576 nats). 4/12 checkpoints are NOISE channels. Upper edge survives the metric change; doc 22 §3 family hypothesis rejected
A19 — re-read the granularity dividend A18's A_eff is alphabet-size invariant once divided by ln m, so doc 16's m=2/4/8/16 ladder became comparable for the first time — at zero model cost Ran; see 26. Dividend holds (I(Z;A) up 3/3, m=16 best 3/3) but is bounded by the encoder, not the alphabet — effective range 2.02–4.98 against a 16-level offer. Doc 16's "no model uses more than 10" overstated it 2–3×; its non-monotonicity is in η, not the encoder
A20 — the ladder on a non-qwen family Doc 26 §5 named "one family" as its main limit: all three granularity rungs are qwen. Does the dividend hold in a family with a genuinely strong verbal encoder? Ran (gemma2:2b); see 27. R1 falsified — peaks at m=4, collapses 9× by m=16, corr −0.46. The dividend is 3/4 and qwen-specific. gemma2 does populate the alphabet (A_eff 4.27) and gains nothing (η 0.21 → 0.01), so doc 26's "wasted, not harmful" is family-specific
A21 — the ladder on mistral and llama Doc 27 left the dividend at 3/4. mistral:7b (NOISE at m=4) and llama3.1:8b (LIVE) have near-identical A_eff and opposite channel status — does granularity act on range, or on whatever makes a channel live? Ran; see 28. A dead channel is not rescued — mistral 0.0000 at all four m (p 0.71–1.00) while A_eff grows 1.08 → 2.72, and its ΔG_self degrades monotonically. llama peaks at m=8. Optimal alphabet is per-checkpoint (16/16/16/8/4/none). Doc 01 §3ii strict version fails 0/6
A22 — the ladder on phi3 phi3 is the last of five families without a granularity ladder, and phi3:medium is NOISE at m=4 like mistral:7b — does doc 28's dead-channel result follow the class or the family? Declined, not run; see 29a. The mandatory parse check failed at a 32-token cap for reasons rule 6 does not cover: phi3 echoes its answer (parser takes it) and overshoots the range (rates "9" on a 1–8 scale). Joint-parse subset ~26–40/80 and non-random. Second time phi3 has been excluded by the shared parser — coverage in this repo is not family-neutral
A23 — bootstrap CIs and a noise floor for I(Z;A) Doc 27 §4 found 2 items of 80 moving an MI ~40%. Which published orderings survive, and what MI does pure noise produce at this n? Ran; see 29. Floor quadruples with m (0.018→0.081). 0/6 claimed optima separate, so doc 28's headline is withdrawn and only 2/6 clear the m=16 floor. Doc 25's LIVE class is independently confirmed — identical to the above-floor set. Negative claims survive, magnitude claims mostly do not
A24 — rerun the ladder at n=200 Doc 29 §6: "the honest fix for most of what A23 withdrew is more items, not better statistics." Does 2.5× the data resolve the optima? Ran; see 30. No — still 0/6 separate, and 2/6 flipped (qwen2.5:14b m=16→m=4). The floor falls 2.6× so detectability improves, but no ordering is resolved. gemma's m=4 fell 29% as doc 27 §4 predicted; its m=16 became significant, so "collapse to zero" was partly power. Encoder gap 6/6
A25 — rerun A18 at n=200 Doc 29 §4 could not compare the twelve-checkpoint table against an n=200 floor, because a18.json held n=80 only. Does the saturation structure survive more data? Ran (12/12); see 31. A_eff stable (mean Δ 0.08), MI fell in 8/9 — n=80 MI is systematically optimistic. Upper edge survives a third test. qwen2.5:1.5b leaves the DEAD class, and doc 29 §4's LIVE≡above-floor coincidence does not replicate
A26 — the m=8/16 rungs at n=500 Doc 30 §6: those rungs were judged where the floor was doing most of the work. Are the "below floor" channels absent, or just under-powered? Ran; see 32. Under-powered — both undecided cases are real at n=500 (p=0.0016, p<10⁻⁴); m=16 above floor goes 2/6 → 4/5. MI rose in 7/10, so its small-n bias direction depends on alphabet sparsity: small alphabets over-read, large ones under-read. First CI-separated ordering in the repo
A5b — adaptation without a running total A5 §4: the prompt displayed cumulative points, confounding rate with accumulation Ran; see 34. The running total was not the cause (slopes barely move) — but 35 shows the window was, so A5b controlled the wrong variable and its "genuinely absent" headline is retracted. Its B3 asymmetry stands

A1–A3 are the load-bearing ones. A4 is the easiest falsification in the whole repo and should be run first for exactly that reason. A7 was the cheapest remaining test — the first here to predict a dissociation within a single model (report channel dead, behaviour channel live) rather than a difference across the ladder, so it could not be explained away as a capability confound. It ran and did not confirm: the behavioural probe chosen (resample stability) turned out not to be independent of the self-model, so it inherited the bound it was meant to escape (11). A7b inherits the role, with an intervention-based probe that genuinely sits outside the self-model.

A7b then also failed (12), which moves the scope limit from untested to unsupported. A7c inherits the role, and it is a change of scale rather than of probe.

Eight methodological rules, each earned by a failed test:

  1. (doc 11 §6, amended by 13 §4) Before claiming a channel is independent of an agent's self-model, check whether it is downstream of the same internal variable — and then check whether they share an encoder. Two observables of one state are two channels when their encoders differ, which is exactly how resampling beats self-report at 14B.
  2. (doc 12 §2) Before counting a behavioural feature as a reading channel, check whether constructing or scoring it requires the ground truth you are trying to infer. A probe built from gold is no evidence for an observer who lacks gold.
  3. (doc 15 §5) Before concluding an agent lacks an internal signal, vary the task format. Every null in docs 0714 was measured on MCQA, the least favourable format; multiple-choice manufactures confident errors and so understates self-knowledge.
  4. (doc 21 §1) A verdict that turns on which checkpoints happened to be installed is not a verdict. Doc 18's "detector, not a gauge" reversed at n = 5 (doc 20); doc 21's own U1 falsification reversed once the other three checkpoints came back. Both times the sub-sample was chosen by convenience, not design. Report coverage as a limit, and re-run before generalising across it.
  5. (doc 22 §1) When you correct a statistic, re-derive every row — including the ones you believe are unaffected. Doc 19 fixed M2 from share-at-level-4 to modal share, then asserted the qwen rungs were exempt because their mode is the top level. It was true for two of three: the qwen-3B's mode is level 1, and its stale 0.34 survived the very fix that should have caught it. An exemption asserted rather than computed is how an error outlives its own correction.
  6. (doc 23 §1) A harness parameter tuned on one family is a filter, not a constant. A 6-token cap on confidence replies was invisible for models that answer "4" and silently truncated every model that preambles — phi3 was scored "insufficient data" at 5/12 when it is 12/12 at 32 tokens. Errors 1–5 mis-scored models being measured; this one excluded a family from measurement, by verbosity style, which correlates with lineage. Before adding a family, verify every fixed budget is non-binding for it.
  7. (doc 24 §1) Never cache a null result. A cold start returned HTTP 200 with an empty body; the empty was cached; a cached empty replays forever. gemma2:2b and qwen2.5:1.5b were both written off as unmeasurable by a transient blip frozen into a permanent verdict. Retry once on empty, and if it is still empty, surface it rather than persisting it.
  8. (doc 27 §1) One estimator per quantity, repo-wide. A18/A19 reimplemented MI as a sum of separately corrected entropies — equivalent to the standard correction only when every joint cell is occupied, and under-correcting increasingly with alphabet size, i.e. toward the granularity dividend it was testing. The tell was there: the new column silently disagreed with docs 15/16/24 for the same checkpoints. If a quantity has an implementation, import it; if a new one is truly needed, diff it against the old on real data before publishing.

Status (2026-07-29): the pilot ransim/real/stage1.py, Qwen2.5 3B/7B, ~1,000 calls, write-up in 07. Scoreboard: the alexithymia bound and capability-tracking supported; appetitive-arm dissociation supported (7B); dissipative arm confounded (needs the behavioural hint-following assay); granularity dividend untestable at this scale (bottleneck upstream of the report alphabet — needs a taller ladder); A4 falsified as operationalized — see the amended reading below.

The A4 arc closed the same day — three runs, ending in a full retraction. 08 (rescaled numerals) corrected doc 07's reading; 09 (mixed units, packs) then dissolved both: with display and value decoupled, λ has no measurable effect on either model (7B: λ=1 vs λ=5 at identical display → spend 0.582 vs 0.582; 3B fully stakes-blind). The earlier "falsification" of monotone arousal was a numeral artifact. A4 final status: inconclusive — 3B–7B subjects never engage the value-state; stated stakes and budgets are cheap talk at this scale. The arousal row is relabeled untested, and the next A4 must use real budgets (enforced token/compute rationing) and real stakes (consequences the run actually bears), on frontier-scale subjects. Then: price-vector arousal (Stage 3 item 1); a 10-model ladder for A1/A3, mirroring the parent repo's v2. Doc 07's L1/L2 results are untouched by the retraction (nothing there was asserted-in-text).

Stage 2 — Human data (reanalysis before collection)

Every claim below can be tested against existing public datasets; none needs new subjects to start.

  • Bivariate valence (doc 02). Reanalyze affect-report corpora asking whether the two arms separate once causal attribution is controlled. The theory predicts a specific 2-factor structure with a predicted loading pattern, not just 2 factors.
  • χ = K·H(k̂) interaction (doc 04). Goal-conflict inventories crossed with stakes measures: the theory predicts a multiplicative, not additive, interaction. This is a clean, unusual, pre-registerable prediction.
  • Adaptation asymmetry (doc 06 §4). Does arousal adapt less than valence? The theory says it should. Longitudinal wellbeing panels with separate arousal items can answer this from existing data.
  • Yerkes–Dodson interaction (doc 03 §2). The peak-shift with task uncertainty is already in the literature; the test is whether the quantitative shift matches argmax_α Σ q log(b_α/r) under measured belief noise.

Stage 3 — The honest negatives to look for

The parent repo's credibility comes substantially from reporting its failures (value/06 R5 "no demon on correlated agents"; the doc-04 sum-form erratum; the doc-01 over-determination retraction). The affect layer already has one such correction (doc 06 §2, where the simulation refuted the stated boundary and produced a better law). Others to hunt for deliberately:

  1. Does λ = K/E collapse stress-from-ambition and stress-from-depletion when the data say they differ? Doc 03 §1 makes this unavoidable. If they dissociate behaviourally, the scalar-λ model is dead and arousal needs the price vector π.
  2. Can any self-report instrument separate v⁺ from v⁻? If not, doc 02's central claim is true-but-untestable by report and needs a physiological assay — which should be said plainly rather than worked around.
  3. Does the K·H(k̂) form survive goal restructuring? Doc 04 §4 concedes the entropy is incomparable across a changing channel set, and the interesting cases of conflict are exactly those. This may be a fatal scope limit rather than a caveat.

Non-goals

  • Any claim about phenomenal experience. Permanently out of scope (doc 00 §3, §5).
  • A single scalar "total emotion." The parent repo's §0 impossibility applies unchanged: value is frame-relative, so affect is too, and there is no god's-eye affect sum. Coordination goes through λ (doc 03 §3), never through addition.
  • Clinical application. Docs 04 and 05 touch depression and akrasia structurally; nothing here is a claim about treatment, and the narrowing result in 04 §3 is explicitly not advice.