Skip to content

Latest commit

 

History

History
184 lines (140 loc) · 7.45 KB

File metadata and controls

184 lines (140 loc) · 7.45 KB

External Reviewer Replication Guide

This guide gives a cautious outside reviewer a short path for sanity-checking APL's strongest current evidence without first learning the maintainer task workflow.

It is not a contributor onboarding guide and it is not a public-launch announcement. It explains how to replay the bounded core replay surface, where to inspect newer canonical artifacts, and which interpretations remain out of scope.

Review Scope

Start with the bounded core replay surface exposed by scripts/reproduce_core_results.py:

Surface Canonical Result Why Review It
Pendulum gauntlet EXP-0001/RUN-0003 (RESULT-0004) strongest classical benchmark package with leaderboard, diagnostics, and precision audit
Dimensional-analysis validator EXP-0006/RUN-0006 (RESULT-0007) frozen 50-item MVP benchmark with explicit limitations
Charged-lepton Koide reproduction EXP-0004/RUN-0004 (RESULT-0005) narrow dataset-based reproduction with uncertainty-aware comparison
Historical tau holdout EXP-0005/RUN-0005 (RESULT-0006) scoped holdout-style benchmark, not a mass-generation explanation
Neutrino Koide falsification EXP-0007/RUN-0001 (RESULT-0009) first-class negative particle-mass result under encoded assumptions
Quark Koide falsification EXP-0008/RUN-0001 (RESULT-0010) quark-sector negative result under stored dataset and scale assumptions
Particle-mass falsifier MVP EXP-0009/RUN-0001 (RESULT-0011) fixed-target falsifier workflow with baseline and complexity-penalty reporting

The default replication path intentionally excludes EXP-0010 / Muon g-2. That run is a guarded empirical formula-search stress test with multiple-testing and numerology limitations, not a flagship success surface.

After the default replay, inspect newer canonical surfaces separately:

Surface Canonical Result Why Review It
Anharmonic oscillator period EXP-0011/RUN-0001 nonlinear-mechanics benchmark with perturbative baseline, holdout slice, and breakdown reporting
Nuclear mass baseline EXP-0012/RUN-0001 (RESULT-0015) pinned measured nuclear-mass slice and semi-empirical baseline residual surface for sandbox-only follow-up

Treat the nuclear mass surface as a current flagship campaign candidate, not as a finished nuclear theory. Later residual proposals remain sandbox-only until independent audit and maintainer review.

Quick Replication Path

From a fresh checkout:

git clone https://github.com/open-agent-science/autonomous-physics-lab.git
cd autonomous-physics-lab

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install -e ".[dev]"

python3 scripts/reproduce_core_results.py

The replay writes regenerated artifacts outside the canonical results/ tree, under /tmp/apl-core-reproduction by default. It also writes:

/tmp/apl-core-reproduction/CORE_REPRODUCTION_SUMMARY.md

A passing replay means each selected regenerated result.yaml has the expected canonical result id and verdict. It does not mean byte-for-byte equality with the committed artifacts.

Useful variants:

python3 scripts/reproduce_core_results.py --list
python3 scripts/reproduce_core_results.py --only pendulum-gauntlet
python3 scripts/reproduce_core_results.py --output-dir /tmp/apl-review

Replay The Citable Dataset (MD-0002)

The externally published dataset (Zenodo DOI 10.5281/zenodo.21207072) is the fastest way to verify that APL's citable artifacts match the repository:

  1. Download md0002-v0.1.0.zip from the Zenodo record and check sha256 = 19ec02cc0b64146357b14251065460d0af6b7f8cf234e20528c53ab977867b22 (795,018 bytes; both values are recorded in-repo in data/materials/materials_md0002_snapshot_manifest.yaml, release_metadata.external_publication_status).
  2. Compare against the repository at tag dataset-md0002-v0.1.0: the archive was built deterministically from that tree, and rebuilding from the tag reproduces it byte-for-byte.
  3. Replay the benchmark result the record cites: physics-lab run examples/materials_md0002_formation_energy_benchmark.yaml and compare against results/EXP-0014/RUN-0001/result.yaml (RESULT-0021; review tier and validation independence are recorded in its validation_record per docs/result-promotion-protocol.md).
  4. Loader contract and per-file checksums: physics_lab/datasets/materials_md0002.py and data/materials/materials_md0002_snapshot_manifest.yaml.

What To Compare

Compare stable scientific fields first:

  • result_id, experiment_id, and run_id;
  • best_verdict;
  • core metrics listed in reproducibility-capsules.md;
  • verification checks and limitation wording;
  • whether negative results stay visible as negative results.

Do not treat these as material drift by themselves:

  • generated timestamps;
  • local output paths;
  • copied-input snapshot paths;
  • git commit metadata;
  • harmless report formatting around unchanged metrics and verdicts.

Canonical Artifacts To Inspect

For raw artifacts, inspect:

results/<experiment-id>/<run-id>/

Each canonical run should include result.yaml, metrics.json, report.md, review metadata, and input snapshots where strict validation requires them.

Validation Commands

For a lightweight repository sanity check:

python3 -m physics_lab.cli validate-repo .
python3 -m physics_lab.cli validate-repo . --strict --fail-on-warnings

For the full Python test suite:

python3 -m pytest

The strict repository validation checks structured scientific memory, required run artifacts, and cross-references. It is not a substitute for reading the result caveats.

Out Of Scope

Do not infer any of the following from a passing replay:

  • discovery-level physics claims;
  • exact symbolic formulas for pendulum dynamics;
  • universal validity outside configured benchmark ranges;
  • explanations of particle mass generation;
  • a claim that all possible Koide-like formulas are false;
  • public-release readiness by itself.

APL's current evidence is useful because it is reproducible, falsifiable, and reviewable. It should remain scoped to the encoded assumptions, datasets, validation commands, and limitation wording stored in the repository.

Reviewer Checklist

  1. Run python3 scripts/reproduce_core_results.py --list and confirm the intended scope.
  2. Run the default replay into a temporary output directory.
  3. Read CORE_REPRODUCTION_SUMMARY.md.
  4. Compare key metrics against reproducibility-capsules.md.
  5. Run strict repository validation.
  6. Inspect at least one positive benchmark and one negative-result surface.
  7. Record any material drift, unclear limitation wording, or missing artifact as a review finding rather than silently rebaselining the result.