ABI owns acquisition and packaging of capabilities from foreign open-weight teachers: source qualification, probing, semantic labeling, quarantine, normalization, provenance, information accounting, minimization research, immutable artifacts, and fail-closed verification.
LayerCake owns runtime hosting, cake installation, transfer between compatible LayerCake hosts, routing, orchestration, and inference performance. Keep the repositories and evidence lineages separate. ABI may link to LayerCake and test its public host interface; it must not copy LayerCake implementation or represent LayerCake research as ABI evidence.
The controlling technical release is R7:
- public tag:
abi-final-validation-v2-repaired-r7-2026-08-30 - tagged commit:
3f82a9f4d67dda5c8ea13bd59b2d8f1bbd3dd128 - release: https://github.com/Yoder23/abi/releases/tag/abi-final-validation-v2-repaired-r7-2026-08-30
- archive SHA-256:
fc50f423986149b5d4670ec9e28698540f64be96034efa26e5704c4469921e88 - strict certificate:
results/abi_final_validation_v2/strict_validation_r7_tar_bound.json - blind review:
results/abi_final_validation_v2/blind_redteam_r7/
R7 passes bounded technical validation. Human review, different-hardware reproduction, and registered minimum-information certification are now open, not complete.
The additive R8 native-neural-transfer campaign is a Level 0 negative result, not a new release line. Its v10 package-only hypernetwork failed the public Pythia prerequisite: AFTER−BASE was +0.001953 across 1,024 paired rows (95% bootstrap CI -0.019531 to +0.024414). The held-out commitment was never revealed. R8 therefore does not establish native cross-model transfer and does not alter R7.
The additive R9 neural-ISA diagnostic is also negative and does not alter R7. Its stricter v2 capability-specific Pythia backend achieved 0.128472 training fit and 0.125 unseen-depth AFTER accuracy versus 0.073242 BASE; ZERO also reached 0.125. Exact live replay covered 8,192 evaluation and 288 training rows, and 5/5 hostile controls rejected tampering. The recipient-state GRU branch is closed; the capability-blind universal backend was not run.
The additive R10 runtime copy/paste campaign separates a bounded component pass from an overall negative result. Identical 2,040-byte source-extracted canonical-IR packages achieved 1.0 AFTER and RESTORED accuracy on all 61,440 Pythia, Qwen2, and T5 recipient rows, with exact removal and exact fresh live replay. The source model's native decoder achieved only 0.59375 to 0.65234375 on the newly sampled composed prompts, below the registered 0.99 source gate. R10 therefore proves only synthetic runtime-owned package paste/remove/restore; it does not prove lossless source behavior, LayerCake integration, native neural transfer, or teacher extraction, and it does not alter R7.
The four published immutable capability packages execute through the canonical capability runtime/conformance interface in the declared LayerCake v25, Qwen2.5-0.5B, and Pythia-160M environments. The proof includes:
- three capability-blind physical certification environments;
- 301,543 content-bound reachable-filesystem rows;
- 5,043 locked matrix rows;
- 3,072 live causality rows in 24 distinct condition processes;
- 2,100 live isolation rows;
- zero capability archives or success IDs in the certification inventories;
- strict recomputation consuming zero scientific status booleans;
- 19/19 fail-closed hostile controls before and after publication;
- exact public-manifest reconstruction in a clean environment; and
- a fresh blind Codex PASS, including 12/12 USTAR/V7 prefix controls.
Never broaden R7 into any of these unproven claims:
- arbitrary-model or universal host compatibility;
- transformer tensor or weight transplantation;
- teacher knowledge extraction, labeling completeness, or source diagnosis;
- teacher-quality English or specialist generation;
- superiority to LoRA, distillation, or fine-tuning;
- human-perceived quality;
- independent different-hardware reproduction;
- a globally minimal English substrate; or
- completion of the full ABI moonshot.
The R7 artifacts are capability-runtime/conformance packages. They are not evidence that ABI has already extracted a compact fluent English core from an arbitrary teacher.
- Preserve every failed release and negative result additively. R5 and R6 are historical failures, not files to rewrite or relabel.
- Treat stored
PASS,status, and gate fields as untrusted. Recompute claims from raw rows, exact counts, hashes, source inventories, and live failure records. - Fail closed on missing packages, raw rows, receipts, source files, hashes, or stale evidence.
- Do not overwrite immutable evidence paths. Use a new revision suffix.
- Do not substitute evidence from a different model, artifact, host, seed, or release lineage.
- Keep human ratings at
0/21,000until three independent raters actually complete and attest the frozen forms. - Keep independent hardware
PENDINGuntil a genuinely independent operator returns the signed different-hardware packet. - Keep minimum-information certification
PENDINGuntil the registered protocol is executed and verified.
- Prefer bounded, preregistered experiments over unbounded nearby sweeps.
- Use GPU for teacher extraction/training when appropriate; CPU-only is not an ABI ideology. Record hardware and one-time/per-host costs honestly.
- Keep generated caches, temporary reconstructions, model weights, and bulk experiment intermediates out of Git. Publish only required immutable assets and verification evidence.
- Run focused tests for changed code, then the full supported suite before a release-ready commit.
- Keep README, active mission, current status, technical claims, final results, review packet, and changelog consistent with the controlling release.
- Execute the frozen human-rating packet with three independent raters.
- Run the external reproduction packet on genuinely different hardware.
- Execute registered minimum-information certification.
- Treat R8, R9, and the overall R10 source-gate failure as preserved falsification results. Any successor native-transfer mechanism requires a new additive preregistration and a materially different canonical IR or recipient injection architecture. It must first fit a capability-specific public control and solve public capability-level realization before universal or held-out use. Do not run further final-state/embedding-state GRU, nearby width, learning-rate, or step-count variants.
- Prove teacher-to-ABI acquisition, capability labeling/segregation, compact English extraction, and quality comparisons against distillation and LoRA only in separately registered lineages.
Do not claim the ABI moonshot until those acquisition and external-quality gates are supported by their own evidence.