Skip to content

Commit d6b8d2a

Browse files
martex-devclaude
andcommitted
H70: K=2 vindicated on evidence, and the live window says the opposite
Aimed at the deployed spec rather than a new family, because the deployed spec is still validated and still losing money live, and four consecutive family searches had produced one standalone-viable edge and three kills. Ledger 167 -> 170. Hypothesis 11 line 60 reads "Long-only. K=2 FIXED". It was fixed by fiat at the start of the rotation family and inherited untested by every descendant, including the H42b spec running on paper today. H70 varied it and nothing else. The K=2 recomputation reproduces the published deployed figures to the digit -- +42.91% / 1.47 / -29.01% against +42.9% / 1.47 / -29.0% -- so every comparison is against the real incumbent. Gate A fails on the primary K=5 AND on every declared cell: MDD is worse than K=2's at K=3, K=5 and K=8 alike. The verdict therefore does not depend on which cell was primary, which makes this a vindication rather than an inconclusive result. K=2 keeps its place on evidence for the first time. All three pre-registered predictions were wrong. Sharpe runs 1.47 / 1.61 / 1.40 / 1.27 and CAGR 42.91 / 46.23 / 34.53 / 27.13 across K = 2/3/5/8 -- both peak at K=3 -- and MDD never improves on K=2. H66's carry finding does not transfer, and the mechanism is mechanical: carry HARVESTS a premium paid by ~20 near-independent funding streams, so averaging more cuts variance without cutting the mean, while rotation SELECTS and the 4th-8th ranked coins are worse assets rather than more draws of the same edge. This is the first evidence that H65's select-vs-harvest distinction is real, arriving from the opposite direction after H66 withdrew it, and it stays a hypothesis rather than a rule -- one measurement in each of two families is exactly the state that produced the withdrawn refinement. K=3 beats the incumbent on Sharpe (1.61 vs 1.47) and CAGR (+46.23% vs +42.91%) and loses only on drawdown. It is NOT adopted: it was not the declared primary, it fails the registered MDD bar -- the bar meta-finding 8 says decides prop-firm outcomes -- and 0.14 of Sharpe across adjacent K on one path is not a measurement of the K surface. Its numbers go to the owner per the charter rather than being suppressed or acted on. The live window points the other way, and that is the most useful thing here. Replaying every cell over 2026-07-10..2026-08-26 on data/lake-current at L=90, which is what the paper account actually runs, gives K=2 -6.06%, K=3 -7.54%, K=5 -3.85%, K=8 -0.85%, with MDD improving monotonically from -15.48% to -7.09%. So concentration accounts for roughly five of the six points rotation-stop gave up live: the drawdown is not purely bad selection luck, it has a structural component and that component is K. And this does not license changing K. Forty-eight days is 1.7% of the evidence behind the 2,880-day backtest, which says K=8 earns 27%/yr against K=2's 43% with a worse drawdown. Switching because the last seven weeks favoured another setting is textbook recency-chasing. A pre-declared inference is withdrawn in the verdict. Section 6 said a Gate A failure would mean concentration does not explain the live drawdown. It does not follow: the bars are computed on the frozen backtest and cannot answer a question about the live window. The verdict stands on the bars; the further claim does not. The sharpened open question for the H59 divergence hunt is whether the live period is unrepresentative or the K surface has moved, and that needs forward time rather than another slice of the same history. Also honoured here, and overdue: the standing commitment in family-expansion-program.md section 5 to re-validate the deployed book as N grows, last met at 125. rotation-stop reproduces 0.9921 at its original 104 trials against a published 0.992, then scores 0.9889 at 167. rotation reproduces 0.9905 at 65 against 0.990, then 0.9870. Both clear the 0.95 bar; 42 more trials cost 0.003 each. The deployed book is still validated. What it is not is profitable live, which is the separate question H59 opened. data/lake-current refreshed to 2026-08-26. The refresh script touches only the current lake, cannot bump the research epoch, and proves rather than asserts that the frozen lake was left alone: BTCUSDT still 3,249 rows ending 2026-07-09, max |diff| 0.0 across all 3,249 shared days. H70 is registered as time-dependent because its live diagnostic reads the moving lake; its BARS read only the frozen lake and are reproducible. Nothing is deployed, changed, or made paper-eligible. The paper records continue unchanged, which is the correct outcome of a vindication. 619 passed, ruff clean, mypy strict clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent 05d9b1f commit d6b8d2a

10 files changed

Lines changed: 3800 additions & 17 deletions

File tree

PROJECT_MEMORY.md

Lines changed: 47 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@
33
The knowledge file: ledger, results, meta-findings, lessons, open
44
assumptions. PROJECT_STATE.md = what runs now; this = why and what we know.
55

6-
## Trial ledger: 167 registered (166 run, 1 data-blocked: H54). Every new
6+
## Trial ledger: 170 registered (169 run, 1 data-blocked: H54). Every new
77
## spec raises the DSR bar. Do not test without a numbered doc FIRST.
88
##
99
## SOURCE OF TRUTH for the ledger is docs/research/ledger/trials.toml, not
@@ -14,7 +14,7 @@ assumptions. PROJECT_STATE.md = what runs now; this = why and what we know.
1414
## H62 as the carry spec), H64 cointegration KILLED, H65 wide-universe
1515
## carry STANDALONE-VIABLE, H66 cross-sectional carry STANDALONE-VIABLE,
1616
## H67 variance risk premium KILLED, H68 cross-venue dislocation SIGNAL,
17-
## H69 the strategy built on H68 KILLED.
17+
## H69 the strategy built on H68 KILLED, H70 K=2 VINDICATED.
1818

1919
## Hypothesis ledger (docs/hypotheses/, docs/research/)
2020

@@ -182,6 +182,36 @@ assumptions. PROJECT_STATE.md = what runs now; this = why and what we know.
182182
a hypothesis whose returns are concentrated before 2024 should say so
183183
in its verdict.
184184

185+
14. **Breadth feeds edges that HARVEST and starves edges that SELECT —
186+
now with evidence from both directions (H70).** H65 proposed this,
187+
H66 withdrew it, and H70 supplies the missing half. Varying the
188+
deployed rotation's slot count — the one number hypothesis 11 fixed
189+
by fiat ("Long-only. **K=2 FIXED**") and every descendant inherited
190+
untested — gives Sharpe **1.47 / 1.61 / 1.40 / 1.27** at K = 2/3/5/8
191+
and MDD **worse than K=2 at every higher K**. All three
192+
pre-registered predictions (Sharpe rising, CAGR falling, MDD
193+
improving, all monotone) were **wrong**.
194+
**The mechanism is mechanical:** carry harvests a premium paid by ~20
195+
near-independent funding streams, so averaging more cuts variance
196+
without cutting the mean. Rotation *selects*, and the 4th-8th ranked
197+
coins are worse assets rather than additional independent draws of
198+
the same edge. Diluting a selection edge lowers the mean faster than
199+
the variance. **Still a hypothesis, not a rule** — one measurement in
200+
each of two families is precisely the evidential state that produced
201+
the withdrawn refinement last time.
202+
**K=2 is vindicated on evidence for the first time.** K=3 beats it on
203+
return (Sharpe 1.61, CAGR +46.23%) and loses on drawdown; it was not
204+
the declared primary, it fails the registered MDD bar, and it is
205+
**not adopted** — acting on it needs a fresh registration.
206+
**And the live window says the opposite:** over 2026-07-10..08-26 the
207+
same cells give K=2 −6.06%, K=5 −3.85%, K=8 **−0.85%**, with MDD
208+
improving monotonically. So concentration is a **real contributor to
209+
the live drawdown** — about five of the six points — and 48 days is
210+
1.7% of the evidence behind the 2,880-day backtest. Both facts are
211+
true; neither licenses changing K. **The sharpened open question:
212+
is the live period unrepresentative, or has the K surface moved?
213+
That needs forward time, not another slice of the same history.**
214+
185215
13. **A significant spread is not a Sharpe — the info bar has no
186216
variance term (H69).** H68's S2 spread was **+3.17% per 7 days**, CI
187217
excluding zero, breadth 17/20, on 31,752 symbol-days. The strategy
@@ -243,6 +273,21 @@ assumptions. PROJECT_STATE.md = what runs now; this = why and what we know.
243273
in the units the position actually pays in, not the units the
244274
phenomenon is quoted in.**
245275

276+
## DSR re-check at 170 trials (2026-08-28) — the standing commitment, honoured
277+
278+
`family-expansion-program.md` §5 requires re-validating the deployed book
279+
as N grows. Last honoured at 125; run again at 167/170.
280+
281+
| Book | Reproduced at its original N | DSR @167 | Bar 0.95 |
282+
|---|---|---|---|
283+
| rotation-stop (deployed) | 0.9921 vs published 0.992 | **0.9889** | CLEARS |
284+
| rotation | 0.9905 vs published 0.990 | **0.9870** | CLEARS |
285+
286+
Forty-two more trials cost **0.003** each. Third confirmation that the
287+
DSR bar is far less sensitive to ledger growth than was once feared. The
288+
deployed book remains validated; what it is not is profitable live, which
289+
is a different question and the one H59 opened.
290+
246291
## DSR re-check at 125 trials (2026-08-11, scripts/dsr_recheck.py)
247292

248293
Correction candidate 7 CLOSED. Reproduce-first guard passed on both books

PROJECT_STATE.md

Lines changed: 99 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -436,7 +436,7 @@ that moves underneath them cannot serve as their witness.
436436
| Path | Contents | Role |
437437
|---|---|---|
438438
| `data/lake` | through **2026-07-09** | the FROZEN research lake. The input set every published figure was computed on. Immutable until a deliberate, recorded epoch bump. |
439-
| `data/lake-current` | through **2026-08-10** | the CURRENT lake. New research, and the divergence hunt. |
439+
| `data/lake-current` | through **2026-08-26** (refreshed 2026-08-28) | the CURRENT lake. New research, and the divergence hunt. |
440440

441441
The frozen one keeps the plain name because every committed script points at
442442
`data/lake`; renaming would touch the whole corpus and would itself invalidate
@@ -506,6 +506,104 @@ presented as a discovery.
506506

507507
---
508508

509+
## H70 — K=2 VINDICATED, and the deployed book re-validated (2026-08-28)
510+
511+
`OBSERVATION` — two things owed to the **deployed spec** rather than to a
512+
new family. Ledger 167 -> 170.
513+
514+
### 1. The standing re-validation, honoured
515+
516+
`family-expansion-program.md` §5 requires re-validating the deployed book
517+
as N grows. Last honoured at 125.
518+
519+
| Book | Reproduced at its original N | DSR @167 | Bar |
520+
|---|---|---|---|
521+
| **rotation-stop (deployed)** | 0.9921 vs published 0.992 | **0.9889** | CLEARS |
522+
| rotation | 0.9905 vs published 0.990 | **0.9870** | CLEARS |
523+
524+
Forty-two more trials cost 0.003 each. **The deployed book is still
525+
validated.** What it is not is profitable live — a different question.
526+
527+
### 2. H70: was K=2 ever the right number?
528+
529+
`OBSERVATION` — hypothesis 11 line 60 reads *"Long-only. **K=2 FIXED**"*.
530+
It was fixed by fiat and inherited untested by every descendant, including
531+
the H42b spec on paper today. H70 varied it and nothing else. The K=2
532+
recomputation reproduces the published deployed figures exactly
533+
(+42.91% / 1.47 / −29.01% vs +42.9% / 1.47 / −29.0%).
534+
535+
| K | CAGR | Sharpe | MDD | DSR@170 |
536+
|---|---|---|---|---|
537+
| **2 (incumbent)** | +42.91% | 1.47 | **−29.01%** | 0.9994 |
538+
| **3** | **+46.23%** | **1.61** | −32.40% | 0.9998 |
539+
| 5 (primary) | +34.53% | 1.40 | −31.46% | 0.9979 |
540+
| 8 | +27.13% | 1.27 | −31.82% | 0.9931 |
541+
542+
**Gate A fails on the primary and on every cell** — MDD is worse than
543+
K=2's at every higher K — so the verdict does not depend on which cell was
544+
primary. **K=2 is vindicated, on evidence for the first time.**
545+
546+
`OBSERVATION` — all three pre-registered predictions were **wrong**.
547+
Sharpe and CAGR peak at K=3 and fall; MDD never improves.
548+
549+
`INTERPRETATION`**H66's carry finding does not transfer, and the
550+
mechanism is mechanical.** Carry *harvests* a premium paid by ~20
551+
near-independent funding streams, so averaging more cuts variance without
552+
cutting the mean. Rotation *selects*, and the 4th–8th ranked coins are
553+
worse assets, not more draws of the same edge. This is the first evidence
554+
that H65's select-vs-harvest distinction is real — arriving from the
555+
opposite direction after H66 withdrew it — and it stays a **hypothesis,
556+
not a rule**.
557+
558+
### 3. K=3, presented and not acted on
559+
560+
K=3 beats the incumbent on **Sharpe (1.61 vs 1.47) and CAGR (+46.23% vs
561+
+42.91%)**, losing only on drawdown (−32.40% vs −29.01%). It is **not
562+
adopted**: it was not the declared primary, it fails the registered MDD
563+
bar — the bar meta-finding 8 says decides prop-firm outcomes — and 0.14
564+
of Sharpe across adjacent K on one path is not a measurement of the K
565+
surface. **The trade is +3.3pp CAGR and +0.14 Sharpe for 3.4pp more
566+
drawdown.** If wanted, it needs its own registration and a prop-sim
567+
pass-rate comparison.
568+
569+
### 4. The live window says the opposite — the H59 divergence, partly answered
570+
571+
Every cell replayed over the live paper window (2026-07-10 → 2026-08-26,
572+
`data/lake-current`, L=90 as the account actually runs):
573+
574+
| K | live return | MDD |
575+
|---|---|---|
576+
| 2 (incumbent) | **−6.06%** | −15.48% |
577+
| 3 | −7.54% | −15.50% |
578+
| 5 | −3.85% | −11.16% |
579+
| **8** | **−0.85%** | **−7.09%** |
580+
581+
`INTERPRETATION`**concentration accounts for roughly five of the six
582+
points rotation-stop gave up live.** The live drawdown is not purely bad
583+
selection luck; it has a structural component and that component is K.
584+
585+
**And this does not license changing K.** 48 days is 1.7% of the evidence
586+
behind the 2,880-day backtest, which says K=8 earns 27%/yr against K=2's
587+
43% with a worse drawdown. Switching because the last seven weeks favoured
588+
another setting is textbook recency-chasing.
589+
590+
`OBSERVATION`**a pre-declared inference was withdrawn.** §6 said a Gate
591+
A failure would mean concentration does not explain the live drawdown.
592+
That does not follow: the bars are computed on the frozen backtest and
593+
cannot answer a question about the live window. Drafting error, corrected
594+
in the verdict.
595+
596+
**The sharpened open question for the divergence hunt: is the live period
597+
unrepresentative, or has the K surface moved? That needs forward time,
598+
not another slice of the same history.**
599+
600+
`OBSERVATION``data/lake-current` was refreshed to 2026-08-26
601+
(`scripts/refresh_current_lake.py`). It touches only the current lake and
602+
proved rather than asserted that the frozen lake was untouched: BTCUSDT
603+
still 3,249 rows ending 2026-07-09, max |diff| 0.0 on all shared days.
604+
605+
---
606+
509607
## H68 — cross-venue dislocation SIGNALS; H69 shows it is not tradable (2026-08-27)
510608

511609
`OBSERVATION` — pre-registered (commit 3eb0fbc) before any study code

0 commit comments

Comments
 (0)