+{"protocol":{"arms":[{"adaptive_seeds":false,"adversary":false,"arm_id":"B0","custodian":false,"isolates":"the floor, and any imbalance in the bank","iterations":1,"kind":"constant","label":"Oracle-null","memory":false,"model_dependent":false,"preregistered":false,"replication":false,"reviewer":false},{"adaptive_seeds":false,"adversary":false,"arm_id":"B1","custodian":false,"isolates":"the naive baseline everyone actually ships","iterations":1,"kind":"direct","label":"Single-shot","memory":false,"model_dependent":true,"preregistered":false,"replication":false,"reviewer":false},{"adaptive_seeds":false,"adversary":false,"arm_id":"B2","custodian":false,"isolates":"whether iteration alone helps","iterations":3,"kind":"direct","label":"Single-agent + loop","memory":false,"model_dependent":true,"preregistered":false,"replication":false,"reviewer":false},{"adaptive_seeds":false,"adversary":false,"arm_id":"B3","custodian":false,"isolates":"whether role decomposition alone helps","iterations":1,"kind":"uncustodied","label":"Multi-role, no adversary","memory":false,"model_dependent":false,"preregistered":false,"replication":false,"reviewer":false},{"adaptive_seeds":false,"adversary":false,"arm_id":"B4","custodian":true,"isolates":"how much comes from mechanism rather than from agents","iterations":1,"kind":"institutional","label":"B3 + preregistration + custodian","memory":false,"model_dependent":false,"preregistered":true,"replication":false,"reviewer":false},{"adaptive_seeds":false,"adversary":true,"arm_id":"B5","custodian":true,"isolates":"the value of adversarial challenge","iterations":1,"kind":"institutional","label":"B4 + Skeptic","memory":false,"model_dependent":false,"preregistered":true,"replication":false,"reviewer":false},{"adaptive_seeds":false,"adversary":true,"arm_id":"B6","custodian":true,"isolates":"replication, review and memory on top of challenge","iterations":1,"kind":"institutional","label":"Full institution","memory":true,"model_dependent":false,"preregistered":true,"replication":true,"reviewer":true},{"adaptive_seeds":false,"adversary":true,"arm_id":"B7","custodian":true,"isolates":"institutional memory's own contribution","iterations":1,"kind":"institutional","label":"Full - memory","memory":false,"model_dependent":false,"preregistered":true,"replication":true,"reviewer":true},{"adaptive_seeds":true,"adversary":true,"arm_id":"B8","custodian":true,"isolates":"whether spending seeds where they matter beats spending them evenly","iterations":1,"kind":"institutional","label":"Full + adaptive seeding","memory":true,"model_dependent":false,"preregistered":true,"replication":true,"reviewer":true}],"bank":{"items_hash":"d4d1766b0f87c89d326068010926ee00819d5f738e20374119d588f629dacbf1","n_items":60,"truth_lock_hash":"8b459a6cc67b41ae146207362ea65983835a78c8cbf7687df2e93c0ef71051a6"},"claim":"Institutional structure - preregistration, adversarial challenge, independent replication, and evidence-typed memory - improves the accuracy and calibration of autonomous empirical research relative to an unstructured agent, at a measurable cost in compute and tokens.","confidence_as_probability":{"contested":0.3,"speculative":0.4,"suggestive":0.55,"supported":0.75,"well_supported":0.9},"exclusion_rules":["An item whose lifecycle halts before a verdict counts as incorrect for that arm rather than being dropped, because dropping it would reward an arm for failing to answer the questions it finds hard.","An arm whose behaviour is dominated by the language model is reported with its results labelled model-dependent, and is excluded from any claim about mechanism when the run used a mock provider.","Every arm's outcome is reported, including the ones that make the project look bad. There is no rule under which a result is withheld.","The primary comparison is against B0, the arm that answers without looking. B1 was v1's baseline and is model-dependent, which under a mock provider made every comparison in the registered family uninterpretable for mechanism. B1 and B2 are still reported; nothing is compared against them.","Brier score and calibration error are computed only over items where the arm asserted an effect. The confidence rubric measures evidence *for an effect*, so a correct 'no effect' answer necessarily carries weak evidence and scored as gross underconfidence in v1 - an artefact of the mapping rather than a property of the institution. Restricting to assertions is the subpopulation where the rubric's quantity and the scored outcome are the same quantity.","The registered prediction is adjudicated on an interval, not on a point estimate. v1's rule compared two point estimates and returned 'upheld' for a one-item difference on a twenty-item bank.","An abstention is reported, never dropped. 'underpowered' still counts as incorrect for verdict accuracy, because the first exclusion rule refuses to reward an arm for declining the questions it found hardest. It is also counted separately, because a system that knows what it cannot measure is not the same as one that guesses.","Coverage and assertion accuracy are reported together and neither is the primary metric. An arm can drive assertion accuracy to 1.0 by answering only what it is sure of, and coverage is what stops that reading as a good result.","The adjudicated quantity is named in the protocol and the verdict is computed from it. A protocol whose prediction and whose adjudication rule describe different quantities has not registered anything, however precise either one is on its own.","Adaptive seeding may only spend seeds the registration already declared. The full seed set is derived from seed_root at registration and has length max_seeds; escalation chooses how far down that list to go and never which seeds are on it.","The escalation decision reads the development split only. Deciding how much more data to collect by looking at the quantity the verdict will be computed from is optional stopping, and would buy significance rather than resolution.","Only custodied arms are replicated. An uncustodied arm reads the development split, which is fixed by seeds derived from the item id, and returns identical results however often it runs - measured, not assumed: running the ladder twice left B0 through B3 identical to three decimals while every custodied arm moved. Replicating a deterministic arm reports a spread of zero as though it were evidence of stability.","Replicates are averaged within a bank item before arms are compared, so the bootstrap continues to resample items - the population the bank can speak for. Pooling replicates as independent observations would treat three looks at one question as three questions and shrink every interval by a factor the design has not earned."],"metrics":["verdict_accuracy","null_accuracy","brier","expected_calibration_error","false_discovery_rate","usd_per_correct_claim","effect_size_error","coverage","assertion_accuracy"],"prediction":"Replication narrows the ladder rather than reordering it. Averaging three custody draws per arm leaves B8 - B6 on coverage positive with an interval still excluding zero, and leaves B4 - B3 on verdict accuracy spanning zero - the contrast that flipped between the v3 and v4 single draws. If B4 - B3 separates under replication, the v4 reading was right and this protocol's caution was wrong.","primary_metric":"verdict_accuracy","registered_at":"2026-09-01","statistics":{"adaptive_seed_ceiling":24,"adjudicated":{"baseline":"B6","direction":"greater","quantity":"coverage","treatment":"B8"},"adjudication":"named_contrast","alpha":0.05,"baseline_arm":"B0","calibration_scope":"asserted_effects","confidence_order":["contested","speculative","suggestive","supported","well_supported"],"interval":"percentile bootstrap over items","multiplicity":"benjamini-hochberg","pairing":"paired over bank items","replicated_arms":"custodied","replicates":3,"resamples":2000,"verdict_vocabulary":["supported","refuted","no_effect","conditional","inconclusive","underpowered"]},"version":"5"},"protocol_hash":"6bfaa13661c63f3d5aca3c33462998c5807e39cb70cdfd6355f2c3a398eca48c"}
0 commit comments