This guide gives a cautious outside reviewer a short path for sanity-checking APL's strongest current evidence without first learning the maintainer task workflow.
It is not a contributor onboarding guide and it is not a public-launch announcement. It explains how to replay the bounded core replay surface, where to inspect newer canonical artifacts, and which interpretations remain out of scope.
Start with the bounded core replay surface exposed by
scripts/reproduce_core_results.py:
| Surface | Canonical Result | Why Review It |
|---|---|---|
| Pendulum gauntlet | EXP-0001/RUN-0003 (RESULT-0004) |
strongest classical benchmark package with leaderboard, diagnostics, and precision audit |
| Dimensional-analysis validator | EXP-0006/RUN-0006 (RESULT-0007) |
frozen 50-item MVP benchmark with explicit limitations |
| Charged-lepton Koide reproduction | EXP-0004/RUN-0004 (RESULT-0005) |
narrow dataset-based reproduction with uncertainty-aware comparison |
| Historical tau holdout | EXP-0005/RUN-0005 (RESULT-0006) |
scoped holdout-style benchmark, not a mass-generation explanation |
| Neutrino Koide falsification | EXP-0007/RUN-0001 (RESULT-0009) |
first-class negative particle-mass result under encoded assumptions |
| Quark Koide falsification | EXP-0008/RUN-0001 (RESULT-0010) |
quark-sector negative result under stored dataset and scale assumptions |
| Particle-mass falsifier MVP | EXP-0009/RUN-0001 (RESULT-0011) |
fixed-target falsifier workflow with baseline and complexity-penalty reporting |
The default replication path intentionally excludes EXP-0010 / Muon g-2. That
run is a guarded empirical formula-search stress test with multiple-testing and
numerology limitations, not a flagship success surface.
After the default replay, inspect newer canonical surfaces separately:
| Surface | Canonical Result | Why Review It |
|---|---|---|
| Anharmonic oscillator period | EXP-0011/RUN-0001 |
nonlinear-mechanics benchmark with perturbative baseline, holdout slice, and breakdown reporting |
| Nuclear mass baseline | EXP-0012/RUN-0001 (RESULT-0015) |
pinned measured nuclear-mass slice and semi-empirical baseline residual surface for sandbox-only follow-up |
Treat the nuclear mass surface as a current flagship campaign candidate, not as a finished nuclear theory. Later residual proposals remain sandbox-only until independent audit and maintainer review.
From a fresh checkout:
git clone https://github.com/open-agent-science/autonomous-physics-lab.git
cd autonomous-physics-lab
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install -e ".[dev]"
python3 scripts/reproduce_core_results.pyThe replay writes regenerated artifacts outside the canonical results/ tree,
under /tmp/apl-core-reproduction by default. It also writes:
/tmp/apl-core-reproduction/CORE_REPRODUCTION_SUMMARY.md
A passing replay means each selected regenerated result.yaml has the expected
canonical result id and verdict. It does not mean byte-for-byte equality with
the committed artifacts.
Useful variants:
python3 scripts/reproduce_core_results.py --list
python3 scripts/reproduce_core_results.py --only pendulum-gauntlet
python3 scripts/reproduce_core_results.py --output-dir /tmp/apl-reviewThe externally published dataset (Zenodo DOI 10.5281/zenodo.21207072) is the fastest way to verify that APL's citable artifacts match the repository:
- Download
md0002-v0.1.0.zipfrom the Zenodo record and checksha256 = 19ec02cc0b64146357b14251065460d0af6b7f8cf234e20528c53ab977867b22(795,018 bytes; both values are recorded in-repo indata/materials/materials_md0002_snapshot_manifest.yaml,release_metadata.external_publication_status). - Compare against the repository at tag
dataset-md0002-v0.1.0: the archive was built deterministically from that tree, and rebuilding from the tag reproduces it byte-for-byte. - Replay the benchmark result the record cites:
physics-lab run examples/materials_md0002_formation_energy_benchmark.yamland compare againstresults/EXP-0014/RUN-0001/result.yaml(RESULT-0021; review tier and validation independence are recorded in itsvalidation_recordper docs/result-promotion-protocol.md). - Loader contract and per-file checksums:
physics_lab/datasets/materials_md0002.pyanddata/materials/materials_md0002_snapshot_manifest.yaml.
Compare stable scientific fields first:
result_id,experiment_id, andrun_id;best_verdict;- core metrics listed in reproducibility-capsules.md;
- verification checks and limitation wording;
- whether negative results stay visible as negative results.
Do not treat these as material drift by themselves:
- generated timestamps;
- local output paths;
- copied-input snapshot paths;
- git commit metadata;
- harmless report formatting around unchanged metrics and verdicts.
- Project status lists the current major result surface.
- Reproducibility capsules provide replay commands, expected metrics, and caveats.
- Result artifacts index maps canonical result ids to stored run folders.
- Negative results registry keeps falsifications visible alongside successful reproductions.
- Scientific result quality rubric records the maintainer-facing quality assessment for current flagship surfaces.
For raw artifacts, inspect:
results/<experiment-id>/<run-id>/
Each canonical run should include result.yaml, metrics.json, report.md,
review metadata, and input snapshots where strict validation requires them.
For a lightweight repository sanity check:
python3 -m physics_lab.cli validate-repo .
python3 -m physics_lab.cli validate-repo . --strict --fail-on-warningsFor the full Python test suite:
python3 -m pytestThe strict repository validation checks structured scientific memory, required run artifacts, and cross-references. It is not a substitute for reading the result caveats.
Do not infer any of the following from a passing replay:
- discovery-level physics claims;
- exact symbolic formulas for pendulum dynamics;
- universal validity outside configured benchmark ranges;
- explanations of particle mass generation;
- a claim that all possible Koide-like formulas are false;
- public-release readiness by itself.
APL's current evidence is useful because it is reproducible, falsifiable, and reviewable. It should remain scoped to the encoded assumptions, datasets, validation commands, and limitation wording stored in the repository.
- Run
python3 scripts/reproduce_core_results.py --listand confirm the intended scope. - Run the default replay into a temporary output directory.
- Read
CORE_REPRODUCTION_SUMMARY.md. - Compare key metrics against reproducibility-capsules.md.
- Run strict repository validation.
- Inspect at least one positive benchmark and one negative-result surface.
- Record any material drift, unclear limitation wording, or missing artifact as a review finding rather than silently rebaselining the result.