Skip to content

Commit 61139ed

Browse files
martex-devclaude
andcommitted
Separate existence from deployment: add STANDALONE-VIABLE
The standing rule "new features must beat the DEPLOYED system, not zero" answers two questions with one answer: does this edge exist, and should it go in the current book. It is correct for the second and was being applied to the first, so a hypothesis that makes money standalone but less than the incumbent adds was recorded KILLED -- the same terminal status as one with no edge at all. The ledger therefore cannot distinguish H04 mean-reversion ("REJECTED -- decisively", no edge) from H23, killed as redundant, whose input signal measured a CI of [+1.25%, +5.99%]. If the deployed book is ever changed, every edge killed for being redundant TO it is indistinguishable from the edges that were never real. The rule protecting the book from redundancy was also erasing the bench. Adds a third terminal outcome. STANDALONE-VIABLE requires everything a deployment claim requires except the comparison to the incumbent: positive after the full cost model, CI excluding zero, DSR_global >= 0.95, engine-grade, pre-registered. It is not softer -- it is the deployment bar minus exactly one comparison -- and it is not a promotion path. Deployment is unchanged: nothing enters paper or live on standalone merit. Meta-finding 4 stands and both its examples remain correctly not-deployed. No threshold moves, no trial count changes, nothing is re-run. Explicitly no retroactive relabelling. Existing KILLED verdicts stand until re-registered and re-run against the new bar; the graveyard audit's candidates are recommendations, not reclassifications. The honest claim today is that the partition exists, not that it is large. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent ea58971 commit 61139ed

2 files changed

Lines changed: 139 additions & 1 deletion

File tree

‎CLAUDE.md‎

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,16 @@ killed ideas without a new pre-registered spec and a stated reason.
1717
- Kill test (cheap information study) BEFORE any strategy build.
1818
Event-driven engine is the source of truth for strategies; vectorized
1919
screening only pre-engine.
20-
- New features must beat the DEPLOYED system incrementally, not zero.
20+
- New features must beat the DEPLOYED system incrementally, not zero —
21+
for DEPLOYMENT decisions. Amended 2026-08-27
22+
(docs/research/standalone-viable-amendment.md): a hypothesis that
23+
clears the full standalone bar (positive after costs, CI excluding
24+
zero, DSR_global >= 0.95, engine-grade) but does NOT beat the
25+
incumbent is closed STANDALONE-VIABLE, not KILLED. It is not
26+
deployed; it is a live edge on the bench. Existence and deployment
27+
are different questions and the incremental bar only answers the
28+
second. No retroactive relabelling: old KILLED verdicts stand until
29+
re-registered and re-run.
2130
- Paper accounts run only validated/eligible specs; one spec per record
2231
(spec change = archive the record, fresh $5,000 start).
2332
- Live/real-money actions are gated: the runbook
Lines changed: 129 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,129 @@
1+
# Methodological Amendment — separating existence from deployment
2+
3+
Date: 2026-08-27. Status: **DECIDED — owner, 2026-08-27. In force.**
4+
Amends: `CLAUDE.md` standing rule *"New features must beat the DEPLOYED
5+
system incrementally, not zero."*
6+
Prompted by: `docs/research/graveyard-audit.md` §2.
7+
8+
This is a change to what a verdict **means**, not to any threshold. No
9+
existing verdict is reversed by this document, and none may be relabelled
10+
under it without being re-checked against §3.
11+
12+
---
13+
14+
## 1. The defect
15+
16+
The standing rule answers two different questions with one answer:
17+
18+
1. **Existence** — does this edge exist at all?
19+
2. **Deployment** — should this go into the current book?
20+
21+
The rule is correct for (2) and is being applied to (1). A hypothesis that
22+
makes money standalone, but less than the incumbent adds, is recorded as
23+
`KILLED` — the same terminal status as a hypothesis with no edge at all.
24+
25+
`OBSERVATION` — the ledger therefore cannot distinguish between
26+
`docs/hypotheses/04-mean-reversion.md` ("REJECTED — decisively", no edge)
27+
and `docs/hypotheses/23-incremental-features.md` (killed as redundant,
28+
whose input signal H13 measured a CI of **[+1.25%, +5.99%]**).
29+
30+
`INTERPRETATION` — this is a loss of information, and it compounds. If the
31+
deployed book is ever changed or retired, every edge that was killed *for
32+
being redundant to it* is sitting in the graveyard indistinguishable from
33+
the edges that were never real. The rule that protects the book from
34+
redundancy is also erasing the bench.
35+
36+
---
37+
38+
## 2. What is NOT changing
39+
40+
Stated first, because the rule being amended has already earned its keep
41+
and the amendment must not be read as weakening it.
42+
43+
- **Deployment still requires beating the incumbent.** Nothing enters a
44+
paper account, the combined book, or any live path on standalone merit.
45+
The incremental bar is untouched for that decision.
46+
- **Meta-finding 4 stands.** "Info-signal ≠ strategy improvement" was
47+
learned the hard way (7d ranking real at info level, degraded the
48+
walk-forward; shock signal real, fully absorbed by deployed momentum).
49+
Both remain correctly not-deployed.
50+
- **No threshold moves.** `DSR_global ≥ 0.95` is unchanged, per
51+
`mi-trial-accounting-design.md` §4.2.
52+
- **No trial count changes.** These trials are already counted. This
53+
relabels an outcome; it re-runs nothing and spends nothing.
54+
55+
---
56+
57+
## 3. The amendment
58+
59+
A third terminal outcome is added alongside `PASS` and `KILLED`:
60+
61+
> **`STANDALONE-VIABLE`** — this hypothesis cleared a full standalone bar
62+
> on its own merits, and did **not** beat the deployed system. It is not
63+
> deployed. It is recorded as a live edge on the bench, re-examinable
64+
> whenever the deployed book changes.
65+
66+
**The standalone bar is a real bar.** To be recorded `STANDALONE-VIABLE`, a
67+
hypothesis must meet **every** requirement a deployment claim meets, except
68+
the comparison to the incumbent:
69+
70+
1. Positive expectancy **after the full cost model** — fees, half-spread,
71+
participation impact. No gross-of-cost result qualifies.
72+
2. A 95% confidence interval **excluding zero**, by the same block-bootstrap
73+
estimator its family already uses.
74+
3. **`DSR_global ≥ 0.95`** against the global trial count, exactly as a
75+
strategy-grade claim.
76+
4. Engine-grade: produced by the event-driven engine, not a vectorized
77+
screen.
78+
5. A pre-registered hypothesis document, as always.
79+
80+
Anything failing any of these is `KILLED`, as before. There is no partial
81+
credit and no "point estimate was positive" route in.
82+
83+
---
84+
85+
## 4. The risk, and the guard
86+
87+
`INTERPRETATION` — the honest danger is that a softer category rots into a
88+
dumping ground for near-misses, and the ledger quietly stops recording
89+
failure. That would destroy the only asset this project has.
90+
91+
Three guards:
92+
93+
- **The bar in §3 is not softer.** It is the deployment bar minus exactly
94+
one comparison. A `STANDALONE-VIABLE` result is *more* evidenced than
95+
most published retail strategies, not less.
96+
- **No retroactive relabelling from the armchair.** Existing `KILLED`
97+
verdicts stay `KILLED` until re-registered and re-run against §3. The
98+
graveyard audit's candidates (FU-B1, H02) are *recommendations to
99+
re-register*, not reclassifications.
100+
- **`STANDALONE-VIABLE` is not a promotion path.** It cannot become
101+
eligible for paper or live deployment by accumulating time or being
102+
looked at again. It re-enters only through a fresh pre-registration
103+
against whatever the incumbent is on that day.
104+
105+
---
106+
107+
## 5. What this changes in practice
108+
109+
`OBSERVATION` — the ledger's headline currently reads as ~4 survivors from
110+
125 trials. Under this amendment the same history would read as three
111+
groups rather than two: no-edge, real-but-redundant, and deployed.
112+
113+
`INTERPRETATION` — that is a materially different research finding, and the
114+
second group has never been counted. How large it is, is unknown until the
115+
candidates are actually re-run; the graveyard audit identified four
116+
suspects and proved none of them. The honest claim today is that the
117+
partition exists, not that it is large.
118+
119+
---
120+
121+
## 6. Consequence for the near-miss rule
122+
123+
The "near-miss rule" (close a hypothesis that misses its bars, record the
124+
figures) is unchanged in mechanics. Its scope narrows: a hypothesis that
125+
misses **only** the incremental comparison, while clearing §3, is closed as
126+
`STANDALONE-VIABLE` rather than `KILLED`.
127+
128+
A hypothesis that misses any §3 requirement is still closed as `KILLED`,
129+
including one that misses by a hair. Near-miss remains a kill.

0 commit comments

Comments
 (0)