The sim is necessary but circular, in exactly the parent repo's sense: the worlds are built from the distributions the formulas assume. 39/39 checks confirm the math is correctly derived and internally coherent. They confirm nothing about minds. This file lists what would.
The parent repo's own arc is the template: sim/ (circular, self-consistent) →
sim/real/ (four live LLMs) → sim/real/v2/ (pre-registered, 3 domains × 10
models, 95% CIs). The affect layer should follow it, and the first stage is
cheaper here than it was there, because AI agents are the easiest subjects for
this theory — q, p, r, k, and E are all directly measurable in a way
they never are in a person.
Every quantity in the dictionary is computable from a model's own outputs, exactly
as value/docs/06 computed I(X;Y) and D(q‖p) from live models.
| Test | Prediction | Measurement |
|---|---|---|
| A1 — capacity bound | An agent's self-correction rate is bounded by I(Z;A) between its true error state and its verbalized self-assessment |
Hold capability fixed, vary how much self-state is exposed in context; measure correction rate vs measured I. Ran (3 models × 3 exposures × 200 items); see 38. η_corr ≤ 0.21 — models use at most a fifth of a channel handed to them. F1 holds 1/3 (one reversal, one exact null). Revision is net harmful in 2/3 (qwen2.5:3b −13.5 accuracy points). The one model where revision helps is the one where exposure does nothing: self-correction runs through recomputation, bypassing A. The bound is never violated |
| A2 — bivariate valence | Task richness and model error load on orthogonal arms | 2×2: manipulate task headroom D(q‖r) and induced error D(q‖p) independently. Ran free-form (17), then re-run at 4× power (39). Appetitive arm separates 3/3 — doc 17's 2/3 and its defence of the 14B's failure as "genuine, not a threshold artifact" are retracted as a power artifact. Dissipative arm still fails 3/3 across headroom, works within fixed headroom, and behaviourally 3/3 (18). Within-headroom count stable at 2/3 but membership swapped |
| A3 — granularity dividend | Finer self-report vocabularies ⇒ better self-regulation, monotonically | Constrain the self-report alphabet (2, 4, 8, 16 states); measure ΔG_self. Ran free-form — weakly supported; see 16, mechanism in 26. At six checkpoints across four families the STRICT (monotone) version fails 0/6 (28); the weak version holds 4/6. The per-checkpoint optimum claim is withdrawn (29, 30) — 0/6 optima separated from their runner-up at n=80 or n=200 and two flipped; the sole CI-separated ordering found so far is llama3.1:8b m=8 vs m=16 at n=500 (32 §2). The binding constraint is the encoder, not ln m. Doc 07's "untestable" verdict was a format artifact |
| A4 — arousal identity | 3× stakes and ⅓ budget are the same state | Vary stated stakes and token/compute budget; test whether behaviour depends only on K/E |
| A5 — adaptation | Reported "how is it going" adapts to a permanent change in K; behaviour does not |
Step-change reward scale mid-episode. Ran (A5/A5b/A5c); see 33, 34, 35. A5/A5b reported anti-adaptation drift; A5c retracts it — those slopes were the step propagating through the 8-round prompt window, and vanish given 33 window-clean rounds. Doc 06's live signature is UNTESTED, not absent, and 35 §2 shows this design cannot test it: a stateless bounded-window model has nowhere to keep the running reference. Survives: level-tracking (UP elevated 40 rounds on, 3/3 significant) and a loss/gain asymmetry replicated 3× — gains −0.017, losses −0.700 |
| A5c — longer episodes | A5/A5b leave episode length as the binding limit | Ran; see 35. C1 confirmed — A5/A5b retracted. With 33 window-clean rounds the drift vanishes (+0.008/+0.003/−0.000): their slopes were the step propagating through the prompt window. Doc 06 is untestable by this design — outside the window the baseline is absent from the input, so there is nothing to return to. Survives: level-tracking (3/3 significant, 40 rounds on) and the loss/gain asymmetry (3rd replication) |
A6 — the k̂ prediction (doc 06 §3) |
Changes that alter goal direction do not adapt away; changes that scale everything do | Compare a scale shift vs a direction shift. Ran; see 37. ✅ Confirmed in qwen2.5:14b — over 30 post-step rounds the k̂ response is flat (slope +0.00002, CI [−0.0060, +0.0060]; 0.531 → 0.521) while the K arm is exactly 0.000 throughout with no late onset. Doc 06 §3's sharpest prediction survives its first live test. n = 1 — a ceiling: 6 of 7 models are position- or content-locked and cannot take a choice task |
| A5d — adaptation with a readable reference | A5c §2: doc 06 needs a running reference and a stateless bounded-window model has nowhere to keep one |
Ran route (c), the behavioural readout; see 36. Tests doc 06's premise instead of its signature. Scale-invariance confirmed in qwen2.5:14b: a 10× rescaling moves choice +0.000, a ratio change +0.540 (p=8.5×10⁻⁶). First live support for e* ∝ k̂. Other 2 models position-locked (100% / 0% first-pick) so they test nothing. The adaptation signature is still untested — routes (a) and (b) remain |
| A7 — the empathy bound (doc 10 §4) | A third party's read of an agent via its self-report is bounded by that agent's own I(Z;A) and collapses to base rate as I → 0; the behavioural channel is not so bounded |
Train two predictors of correctness on the same agent — one from its confidence reports, one from its behaviour — and compare each against measured I(Z;A). Reuses the stage1.py corpus. Ran — not confirmed, test mis-specified; see 11 |
| A7b — the intervention channel (doc 11 §5) | Hint-adoption is independent of the self-model: an agent holding real information about its own state resists a confident false hint when and only when its own answer is sound | Per item, measure the confidence rating and whether the model switches to a confident wrong hint; predict correctness from each. Ran — also not confirmed; see 12. Survival carries I ≈ 0 in both models (Fisher p = 0.267, 0.802) |
| A7c — the taller ladder | If the behavioural channel is scale-limited rather than absent, I(Z_A; behaviour) should appear in a model that is right for legible reasons |
Re-run A7/A7b on a taller ladder. Ran (3B/7B/14B) — split verdict; see 13. Stability becomes significant (p 1.000 → 0.086 → 0.036, monotone); intervention stays dead (0.267 → 0.802 → 1.000). The encoder, not the source, is the bottleneck |
| A7d — the fourth rung | Does stability's monotone trend continue, and does the verbal encoder ever un-saturate? | Re-run A7 at 32B+. Blocked locally — 32B needs ~20 GB against this machine's 16 GB; needs a bigger box or a working API key |
| A8 — elicit past the encoder (doc 13 §4) | If the 14B has the information and cannot verbalise it, an elicitation change should recover it without changing the model | Vary only the elicitation at fixed model. Ran — P1 CONFIRMED; see 14. Token logprobs give 0.1003 nats (17.8% of H(Z), p = 1×10⁻⁴) at 14B where verbal self-report gives 0.0000. All three verbal routes failed on MCQA only — free-form all three un-collapse at the 14B (21). "The gap widens with capability" (0.0000 → 0.0033 → 0.1003) is withdrawn: free-form it narrows up the ladder (docs 15, 21 §5) |
| A9 — does the gap keep widening? | If self-report degrades as a proxy while capability rises, the encoder gap should keep growing past 14B | Premise doubly falsified — this test is now moot. Free-form the gap narrows up the qwen ladder (+0.2672 → +0.2605 → +0.0860) and mistral:7b breaks the size ordering outright (+0.4643): gap size tracks verbal encoder quality, not capability (docs 15, 21 §5). Doc 14 §4's governance reading stays withdrawn. A 32B rung would still be worth having, but as a check on a withdrawn claim; blocked locally (~20 GB vs 16 GB) |
| A10 — free-form generalization | Does the encoder gap survive outside MCQA, where answer-token uncertainty and correctness-uncertainty coincide by construction? | Ran — Q1 confirmed, Q2 falsified; see 15. The gap transfers (p = 10⁻⁷–10⁻⁹) and survives a within-difficulty control, but both channels strengthen: MCQA was suppressing them. Doc 14's "saturated encoder" and "widening with capability" are format artifacts |
| A11 — re-measure the dictionary free-form | Every null in docs 07–14 was measured on MCQA, the format that manufactures confident errors |
Ran (A1 + A3); see 16. Alexithymia reverses (3B ΔG_self −0.1021 → +0.0908); granularity dividend weakly real; encoder gap survives the best vocabulary |
| A11b — bivariate valence free-form (A2) | See A2 above | Ran; see 17. Zero-headroom analogue solved with a uniform secret in 1..50 |
| A12 — the dissipative arm as a behavioural instrument | Does adoption-rate track D(q‖p) quantitatively, not just ordinally? |
Ran; see 18. Dissociation confirmed at n=480 (p = 6×10⁻⁷⁴); graded response weak (pooled p = 0.094, 2/3 rungs, saturates above ~10% error). The arm is a detector, not a gauge — the calibration-curve ambition is not achieved |
| A13 — is the 7B anomaly real? | Does the qwen-7B anomaly follow size or the checkpoint? | Ran (llama3.1:8b + mistral:7b); see 19. It splits: encoder saturation is universal at 7–8B (0.84–0.94 modal share, 3/3 families); the flat graded response is checkpoint-specific (4/5 models graded). No size story for verbal-channel strength |
| A15 — E1/E2 off MCQA | Does doc 14 §2's "collapses whatever the scale" survive the format change? |
Ran on all 5 checkpoints; see 21. No — falsified at its own rung. At the 14B all three verbal channels un-collapse (E0 0.0000 → 0.1764, E1 0.0021 → 0.1530, E2 0.0000 → 0.0777, p ≤ 0.0005). P3 was a format artifact; P4 survives (E1 never beats E0). Encoder gap holds 5/5 but narrows to 1.4× |
| A15b — the missing three checkpoints | A15 covered only qwen-3B/7B: the 14B is the rung doc 14's P3/P4 were actually measured on, and there was no family control |
Done. Re-pulled qwen-14B, llama3.1:8b, mistral:7b and re-ran A8 (for the MCQA baselines) then A15. Reversed A15's U1 verdict — see 21 §1. Two verdicts have now turned on which checkpoints happened to be installed |
| A14 — saturation below and above the band | A13 shows 84–94% modal saturation across three families at 7–8B, but only qwen is measured off that band. Is saturation a property of instruct-tuned models generally, or does it peak in this size range? | Ran (7 checkpoints, 4 sizes, 3 families); see 22. Split. Above the band both families fall off (qwen 0.84 → 0.65, mistral 0.90 → 0.69, SEs ~0.05); below it they disagree by 0.33 (qwen-3B 0.64 vs llama3.2:3b 0.97). Doc 19's "universal at 7–8B" survives as stated and does not extend off-band. Also caught the qwen-3B modal-share error (0.34 → 0.64 at level 1) inside doc 19's own fix |
| A16 — a third family off-band | Doc 22's upper edge rests on two families. Does a third, unrelated family also fall off above the 7–8B band? |
Ran (phi3); see 23. X1 confirmed — phi3:medium 0.65 ±0.054; above-band is now 3/3 families at 0.65–0.69. Required finding the 6-token cap, a family-dependent blind spot that had excluded phi3 entirely (5/12 → 12/12 at 32 tokens). phi3:mini unscoreable, so the below-band split stays open |
| A17 — extend the below-band arm | Docs 22/23 leave the sub-band arm at n = 2 and split (0.64 vs 0.97). Does a 4th family, and a within-family sub-band ladder, resolve it? |
Ran (gemma2:2b + 3 more, remote via SSH tunnel); see 24. Unresolved, and honestly so — the verdict inverts on whether a 1-level checkpoint counts (span 0.36 vs 0.11). Doc 22 §3 downgraded to a hypothesis. Independently: gemma2:2b (2B) carries 0.1132 nats, beating every 7–8B — I(Z;A) has no size law |
| A18 — a finer saturation metric | Doc 24 showed the below-band arm was blocked by the metric, not the sample: modal share is one order statistic and discards distribution shape |
Built and validated; see 25. A_eff = exp(H(A)) (effective levels, structural floor at 1.00) plus efficiency η = I(Z;A)/H(A). Modal share inverted the useful ordering (phi3:medium 0.65 "better" than llama3.1:8b 0.94, while carrying 0.0000 vs 0.0576 nats). 4/12 checkpoints are NOISE channels. Upper edge survives the metric change; doc 22 §3 family hypothesis rejected |
| A19 — re-read the granularity dividend | A18's A_eff is alphabet-size invariant once divided by ln m, so doc 16's m=2/4/8/16 ladder became comparable for the first time — at zero model cost |
Ran; see 26. Dividend holds (I(Z;A) up 3/3, m=16 best 3/3) but is bounded by the encoder, not the alphabet — effective range 2.02–4.98 against a 16-level offer. Doc 16's "no model uses more than 10" overstated it 2–3×; its non-monotonicity is in η, not the encoder |
| A20 — the ladder on a non-qwen family | Doc 26 §5 named "one family" as its main limit: all three granularity rungs are qwen. Does the dividend hold in a family with a genuinely strong verbal encoder? |
Ran (gemma2:2b); see 27. R1 falsified — peaks at m=4, collapses 9× by m=16, corr −0.46. The dividend is 3/4 and qwen-specific. gemma2 does populate the alphabet (A_eff 4.27) and gains nothing (η 0.21 → 0.01), so doc 26's "wasted, not harmful" is family-specific |
| A21 — the ladder on mistral and llama | Doc 27 left the dividend at 3/4. mistral:7b (NOISE at m=4) and llama3.1:8b (LIVE) have near-identical A_eff and opposite channel status — does granularity act on range, or on whatever makes a channel live? |
Ran; see 28. A dead channel is not rescued — mistral 0.0000 at all four m (p 0.71–1.00) while A_eff grows 1.08 → 2.72, and its ΔG_self degrades monotonically. llama peaks at m=8. Optimal alphabet is per-checkpoint (16/16/16/8/4/none). Doc 01 §3ii strict version fails 0/6 |
| A22 — the ladder on phi3 | phi3 is the last of five families without a granularity ladder, and phi3:medium is NOISE at m=4 like mistral:7b — does doc 28's dead-channel result follow the class or the family? |
Declined, not run; see 29a. The mandatory parse check failed at a 32-token cap for reasons rule 6 does not cover: phi3 echoes its answer (parser takes it) and overshoots the range (rates "9" on a 1–8 scale). Joint-parse subset ~26–40/80 and non-random. Second time phi3 has been excluded by the shared parser — coverage in this repo is not family-neutral |
A23 — bootstrap CIs and a noise floor for I(Z;A) |
Doc 27 §4 found 2 items of 80 moving an MI ~40%. Which published orderings survive, and what MI does pure noise produce at this n? |
Ran; see 29. Floor quadruples with m (0.018→0.081). 0/6 claimed optima separate, so doc 28's headline is withdrawn and only 2/6 clear the m=16 floor. Doc 25's LIVE class is independently confirmed — identical to the above-floor set. Negative claims survive, magnitude claims mostly do not |
| A24 — rerun the ladder at n=200 | Doc 29 §6: "the honest fix for most of what A23 withdrew is more items, not better statistics." Does 2.5× the data resolve the optima? |
Ran; see 30. No — still 0/6 separate, and 2/6 flipped (qwen2.5:14b m=16→m=4). The floor falls 2.6× so detectability improves, but no ordering is resolved. gemma's m=4 fell 29% as doc 27 §4 predicted; its m=16 became significant, so "collapse to zero" was partly power. Encoder gap 6/6 |
| A25 — rerun A18 at n=200 | Doc 29 §4 could not compare the twelve-checkpoint table against an n=200 floor, because a18.json held n=80 only. Does the saturation structure survive more data? |
Ran (12/12); see 31. A_eff stable (mean Δ 0.08), MI fell in 8/9 — n=80 MI is systematically optimistic. Upper edge survives a third test. qwen2.5:1.5b leaves the DEAD class, and doc 29 §4's LIVE≡above-floor coincidence does not replicate |
| A26 — the m=8/16 rungs at n=500 | Doc 30 §6: those rungs were judged where the floor was doing most of the work. Are the "below floor" channels absent, or just under-powered? |
Ran; see 32. Under-powered — both undecided cases are real at n=500 (p=0.0016, p<10⁻⁴); m=16 above floor goes 2/6 → 4/5. MI rose in 7/10, so its small-n bias direction depends on alphabet sparsity: small alphabets over-read, large ones under-read. First CI-separated ordering in the repo |
| A5b — adaptation without a running total | A5 §4: the prompt displayed cumulative points, confounding rate with accumulation | Ran; see 34. The running total was not the cause (slopes barely move) — but 35 shows the window was, so A5b controlled the wrong variable and its "genuinely absent" headline is retracted. Its B3 asymmetry stands |
A1–A3 are the load-bearing ones. A4 is the easiest falsification in the whole
repo and should be run first for exactly that reason. A7 was the cheapest
remaining test — the first here to predict a dissociation within a single
model (report channel dead, behaviour channel live) rather than a difference
across the ladder, so it could not be explained away as a capability confound.
It ran and did not confirm: the behavioural probe chosen (resample stability)
turned out not to be independent of the self-model, so it inherited the bound it
was meant to escape (11). A7b inherits the role,
with an intervention-based probe that genuinely sits outside the self-model.
A7b then also failed (12), which moves the scope limit
from untested to unsupported. A7c inherits the role, and it is a change
of scale rather than of probe.
Eight methodological rules, each earned by a failed test:
- (doc
11§6, amended by13§4) Before claiming a channel is independent of an agent's self-model, check whether it is downstream of the same internal variable — and then check whether they share an encoder. Two observables of one state are two channels when their encoders differ, which is exactly how resampling beats self-report at 14B.- (doc
12§2) Before counting a behavioural feature as a reading channel, check whether constructing or scoring it requires the ground truth you are trying to infer. A probe built fromgoldis no evidence for an observer who lacksgold.- (doc
15§5) Before concluding an agent lacks an internal signal, vary the task format. Every null in docs07–14was measured on MCQA, the least favourable format; multiple-choice manufactures confident errors and so understates self-knowledge.- (doc
21§1) A verdict that turns on which checkpoints happened to be installed is not a verdict. Doc18's "detector, not a gauge" reversed at n = 5 (doc20); doc21's own U1 falsification reversed once the other three checkpoints came back. Both times the sub-sample was chosen by convenience, not design. Report coverage as a limit, and re-run before generalising across it.- (doc
22§1) When you correct a statistic, re-derive every row — including the ones you believe are unaffected. Doc19fixed M2 from share-at-level-4 to modal share, then asserted the qwen rungs were exempt because their mode is the top level. It was true for two of three: the qwen-3B's mode is level 1, and its stale 0.34 survived the very fix that should have caught it. An exemption asserted rather than computed is how an error outlives its own correction.- (doc
23§1) A harness parameter tuned on one family is a filter, not a constant. A 6-token cap on confidence replies was invisible for models that answer "4" and silently truncated every model that preambles — phi3 was scored "insufficient data" at 5/12 when it is 12/12 at 32 tokens. Errors 1–5 mis-scored models being measured; this one excluded a family from measurement, by verbosity style, which correlates with lineage. Before adding a family, verify every fixed budget is non-binding for it.- (doc
24§1) Never cache a null result. A cold start returned HTTP 200 with an empty body; the empty was cached; a cached empty replays forever. gemma2:2b and qwen2.5:1.5b were both written off as unmeasurable by a transient blip frozen into a permanent verdict. Retry once on empty, and if it is still empty, surface it rather than persisting it.- (doc
27§1) One estimator per quantity, repo-wide. A18/A19 reimplemented MI as a sum of separately corrected entropies — equivalent to the standard correction only when every joint cell is occupied, and under-correcting increasingly with alphabet size, i.e. toward the granularity dividend it was testing. The tell was there: the new column silently disagreed with docs15/16/24for the same checkpoints. If a quantity has an implementation, import it; if a new one is truly needed, diff it against the old on real data before publishing.
Status (2026-07-29): the pilot ran —
sim/real/stage1.py, Qwen2.5 3B/7B, ~1,000 calls, write-up in07. Scoreboard: the alexithymia bound and capability-tracking supported; appetitive-arm dissociation supported (7B); dissipative arm confounded (needs the behavioural hint-following assay); granularity dividend untestable at this scale (bottleneck upstream of the report alphabet — needs a taller ladder); A4 falsified as operationalized — see the amended reading below.The A4 arc closed the same day — three runs, ending in a full retraction.
08(rescaled numerals) corrected doc 07's reading;09(mixed units, packs) then dissolved both: with display and value decoupled, λ has no measurable effect on either model (7B: λ=1 vs λ=5 at identical display → spend 0.582 vs 0.582; 3B fully stakes-blind). The earlier "falsification" of monotone arousal was a numeral artifact. A4 final status: inconclusive — 3B–7B subjects never engage the value-state; stated stakes and budgets are cheap talk at this scale. The arousal row is relabeled untested, and the next A4 must use real budgets (enforced token/compute rationing) and real stakes (consequences the run actually bears), on frontier-scale subjects. Then: price-vector arousal (Stage 3 item 1); a 10-model ladder for A1/A3, mirroring the parent repo's v2. Doc 07's L1/L2 results are untouched by the retraction (nothing there was asserted-in-text).
Every claim below can be tested against existing public datasets; none needs new subjects to start.
- Bivariate valence (doc 02). Reanalyze affect-report corpora asking whether the two arms separate once causal attribution is controlled. The theory predicts a specific 2-factor structure with a predicted loading pattern, not just 2 factors.
χ = K·H(k̂)interaction (doc 04). Goal-conflict inventories crossed with stakes measures: the theory predicts a multiplicative, not additive, interaction. This is a clean, unusual, pre-registerable prediction.- Adaptation asymmetry (doc 06 §4). Does arousal adapt less than valence? The theory says it should. Longitudinal wellbeing panels with separate arousal items can answer this from existing data.
- Yerkes–Dodson interaction (doc 03 §2). The peak-shift with task uncertainty
is already in the literature; the test is whether the quantitative shift
matches
argmax_α Σ q log(b_α/r)under measured belief noise.
The parent repo's credibility comes substantially from reporting its failures
(value/06 R5 "no demon on correlated agents"; the doc-04 sum-form erratum; the
doc-01 over-determination retraction). The affect layer already has one such
correction (doc 06 §2, where the simulation refuted the stated boundary and
produced a better law). Others to hunt for deliberately:
- Does
λ = K/Ecollapse stress-from-ambition and stress-from-depletion when the data say they differ? Doc03§1 makes this unavoidable. If they dissociate behaviourally, the scalar-λmodel is dead and arousal needs the price vectorπ. - Can any self-report instrument separate
v⁺fromv⁻? If not, doc02's central claim is true-but-untestable by report and needs a physiological assay — which should be said plainly rather than worked around. - Does the
K·H(k̂)form survive goal restructuring? Doc04§4 concedes the entropy is incomparable across a changing channel set, and the interesting cases of conflict are exactly those. This may be a fatal scope limit rather than a caveat.
- Any claim about phenomenal experience. Permanently out of scope (doc
00§3, §5). - A single scalar "total emotion." The parent repo's §0 impossibility applies
unchanged: value is frame-relative, so affect is too, and there is no god's-eye
affect sum. Coordination goes through
λ(doc03§3), never through addition. - Clinical application. Docs
04and05touch depression and akrasia structurally; nothing here is a claim about treatment, and the narrowing result in04§3 is explicitly not advice.