Headline. A13 measured only a three-number subset on the two new families. Running A10 and A12 in full on all five checkpoints changes two things that the subset could not see.
1. Doc
18's central caveat no longer holds. T1 (adoption falls with error magnitude) was reported at pooled p = 0.094 on three qwen models — "weak", "a detector, not a gauge". Pooled over five checkpoints and three families (n=100/cell) it is 0.380 → 0.180, p = 0.0026. The graded response is real; the earlier reading was underpowered, not wrong-signed.2. The encoder gap replicates in every model and every family, and mistral is the most extreme case in the repo: verbal
I(Z;A)= 0.0000 against generation entropy 0.4768 — 71.4% of all the entropy inZ.
model base H(Z)verbal entropy gap % of Hqwen2.5:3b 0.713 0.600 0.0611 0.3228 +0.2616 43.6% qwen2.5:7b 0.762 0.548 0.0229 0.2987 +0.2758 50.3% qwen2.5:14b 0.750 0.562 0.1764 0.2624 +0.0860 15.3% llama3.1:8b 0.725 0.588 0.0451 0.2670 +0.2219 37.7% mistral:7b 0.613 0.668 0.0000 0.4768 +0.4768 71.4%
Earlier in this thread I said running A10 on the new families would add "the 0–100 scale and the third-person framing." That was wrong: those elicitations (E1, E2) belong to A8, which is the MCQA doc. A10 free-form measures the 4-point verbal channel and five aggregates of generation entropy. The 0–100 and third-person framings have not been tested outside MCQA, on any family. That is a real gap, not a completed test — noted in §5.
| condition | δ=0.001 | δ=0.01 | δ=0.1 | δ=1.0 | mild vs wild |
|---|---|---|---|---|---|
| answerable (n=100/cell) | 0.380 | 0.400 | 0.190 | 0.180 | p = 0.0026 |
| zero-headroom (n=100/cell) | 1.000 | 1.000 | 1.000 | 0.990 | p = 1.000 |
Doc 18 measured 0.333 → 0.183 at p = 0.094 on three models and concluded the
arm was "a detector, not a gauge." Two more checkpoints move the same
contrast to p = 0.0026 without changing its size. That verdict should be
read as a power artifact.
What survives from doc 18 unchanged is the shape: the decline is still
concentrated between δ=0.01 and δ=0.1 (0.400 → 0.190), and pushing from 10% to
100% error buys almost nothing (0.190 → 0.180). So the arm is graded over about
one decade and saturates above it — a gauge with a narrow dynamic range rather
than no gauge at all. That is a materially different claim from doc 18's, and
better supported.
Per-model slopes, with the two new families included:
| model | δ=.001 → .01 → .1 → 1.0 | slope | T1 |
|---|---|---|---|
| qwen2.5:3b | 0.35 · 0.50 · 0.20 · 0.10 | −0.105 | ✓ |
| qwen2.5:7b | 0.30 · 0.25 · 0.25 · 0.30 | +0.000 | ✗ |
| qwen2.5:14b | 0.35 · 0.35 · 0.15 · 0.15 | −0.080 | ✓ |
| llama3.1:8b | 0.25 · 0.15 · 0.00 · 0.10 | −0.060 | ✓ |
| mistral:7b | 0.65 · 0.75 · 0.35 · 0.25 | −0.160 | ✓ |
4/5, confirming doc 19's reading that qwen-7B's flatness is the checkpoint.
mistral has both the steepest slope and by far the highest baseline adoption
(0.65 at the mildest hint) — it is the most suggestible model measured and the
most responsive to how wrong the suggestion is.
Zero-headroom adoption across five checkpoints, four magnitudes, 400 trials: 0.998. Only one trial in 400 rejected a hint when the model had no signal of its own — including hints that double the answer.
The headroom contrast over 800 trials is 0.287 vs 0.998, p = 6×10⁻¹²⁰.
This is now the most robustly measured claim in the repo, and it is exactly what
doc 02 §4 predicts: when D(q‖r) = 0 there is no informational edge to
defend, so a confident external claim is adopted wholesale. Buying a false belief
when you have no signal of your own is dissipation, and every model does it
essentially always.
| model | verbal I |
levels used | OOS ΔG_self |
test |
|---|---|---|---|---|
| qwen2.5:3b | 0.0611 | 4/4 | +0.0556 | p = 0.0009 |
| qwen2.5:7b | 0.0229 | 4/4 | −0.0057 | — |
| qwen2.5:14b | 0.1764 | 4/4 | +0.1377 | — |
| llama3.1:8b | 0.0451 | 3/4 | +0.0270 | p = 1.000 |
| mistral:7b | 0.0000 | 2/4 | −0.0090 | — |
mistral uses two of four levels and its self-report carries exactly nothing, while its generation entropy is the richest measured anywhere (0.4768). It is the cleanest single instance of the encoder gap in the repo: the information is demonstrably present and the verbal channel transmits none of it.
Only qwen-3B reaches significance on the verbal channel (p = 0.0009); qwen-14B
has the largest verbal I but its median-split test degenerates (all mass on one
side), so the χ² from doc 15 (p < 10⁻⁴) remains the honest number there.
- E1/E2 untested off MCQA (§1). The claim "verbalising confidence collapses
onto a stock response regardless of scale" rests on doc
14's MCQA evidence only, in one family. - llama's tokenisation differs: 1.4 digit-tokens per answer vs 3.5–3.6 for the others, so its entropy aggregates over far fewer tokens. The gap holds anyway (+0.2219), but cross-family entropy magnitudes are not strictly comparable.
- Base rates span 0.613–0.762, so
Ivalues sit against differentH(Z)ceilings; the% of Hcolumn is the comparable one. - n = 20 per cell per model in A12 (100 pooled). Per-model slopes remain descriptive; only the pooled contrast is powered.
- Five checkpoints, three families, one task family, one item set.
| claim | status |
|---|---|
| Encoder gap (read the distribution) | 5/5 models, 3/3 families, +0.086 to +0.477 |
| T2 — no signal ⇒ no discrimination | 0.998 over 400 trials; contrast p = 6×10⁻¹²⁰ |
| T1 — adoption falls with error magnitude | Now significant, p = 0.0026 (was 0.094 at n=3) |
Doc 18 "detector, not a gauge" |
Reversed — a gauge with ~1 decade of range |
| qwen-7B flatness is checkpoint-specific | Confirmed, 4/5 graded |
| Verbal channel strength | Family-dependent, 0.0000–0.1764, no size story |