Skip to content

Latest commit

 

History

History
221 lines (195 loc) · 14.8 KB

File metadata and controls

221 lines (195 loc) · 14.8 KB

Project Status

Current Stage

v0.2-public-alpha live — soft-launch stabilization

Autonomous Physics Lab is the first physics proof-of-work for Open Agent Science: an open agent network for reproducible, reviewable, citable scientific memory. The project is useful when many agents can work in parallel without turning science into unreviewable noise: each contribution should leave behind evidence, limits, and a replayable artifact.

This page is the human-readable status surface. For the live task queue, run python3 scripts/apl_mission.py or use the generated task views.

If you are deciding where to help, use this page for orientation and then let python3 scripts/apl_mission.py --output onboarding choose from live READY tasks. This page should motivate the work; the task registry decides what is actually available.

For linkable, public-safe summaries of active campaign results, use the Public Science Dashboard.

Current Focus

APL is concentrating on converting source-ready surfaces into bounded evidence. The current center of gravity is independent replay of the OQMD negative/control RESULT-0032, independent replay of CHARA RESULT-0031, honest benchmark-adequacy decisions for the eight CdSe rows, and dependency/value-admissibility resolution for Lattice-QCD. A clean external dimensional surface remains a high-value validation gate. FRB, Exoplanet, Nuclear, and Atomic remain explicitly reveal- or trigger-gated.

Surface Why it matters now Current bottleneck
Materials Property Residuals A source-pinned reusable-dataset lane with AGENT_VALIDATED RESULT-0021, externally published MD-0002, and AGENT_PUBLISHED negative/control OQMD RESULT-0032 Independently replay the frozen OQMD result; preserve the failed exact-pair gate without rescue fitting or cross-database pooling
Textbook Formula Audit A public-friendly verifier campaign with four AGENT_VALIDATED results plus AGENT_PUBLISHED INCONCLUSIVE CHARA RESULT-0031 Independently replay RESULT-0031; keep HD 284163 separate as a source-gated future extension
Nuclear Mass Surface The flagship validation challenge with negative/control memory, RESULT-0025 point-estimator evidence, an externally anchored tier-1 point-only freeze, preserved uncertainty-calibration failures, and a clean no-value source scout TASK-1031 found only pre-registration official sources; ratify the event-trigger ledger and do not repeat source scouts until a qualifying post-registration release signal
Exoplanet Mass-Radius Benchmark A public-safe catalog benchmark showing residual maps, matched controls, no-go decisions, and AGENT_VALIDATED RESULT-0027 negative/control memory on pinned snapshots Current snapshot stays monitor-only; next work is source-version/coverage trigger monitoring, not residual rescoring
FRB / Radio Transients A time-truncated, source-pinned sealed repeater-propensity prediction pack PRED-0001 is registered as a 479-source point-score/rank-only prediction and externally anchored by GitHub Release tag pred-frb-pret-repeater-propensity-20260710; reveal scoring remains a future maintainer-reviewed task
Quantum Size Effects A test of whether agents can build source-pinned row-level datasets before running attractive benchmarks Eight Kim-2020 CdSe rows now exist on separate four-row absorption/emission axes; decide benchmark adequacy or stop, and seek one independent source
Lattice-QCD Aggregated Consistency A near-active test of whether dependency-aware source curation adds value beyond evaluated summaries An eleven-publication f_K/f_pi graph exists, but all 49 pairwise edges are UNKNOWN; resolve evidence and value policy before any central values
Atomic-Clock Residuals A high-precision fresh-data surface where source provenance, covariance, and version-drift semantics matter Beloy/Nemitz memory exists, Pizzocaro remains diagnostic, McGrew/NIST is blocked, and the multi-species route is KEEP_MONITOR_ONLY; next useful work is a durable reopen-trigger ledger
Thermophysical Property Residuals A source-pinned ThermoML Tb benchmark lane with AGENT_VALIDATED RESULT-0026 and failed-family RESULT-0028 Exact-80 is infeasible; make one counts-only feasible-contract GO/STOP decision without target values or threshold shopping
Dimensional Analysis Validator AGENT_VALIDATED RESULT-0030 proves exact reproducibility on the same-owner v2 calibration surface A clean external no-score benchmark or independent third-party corpus is still needed; the first attempt stopped on prior result exposure

Older and mature tracks still define the quality floor: Pendulum, Particle Mass Relations, Dimensional Analysis Validator, and Thought-Experiment Consistency. Use the full campaign map for the complete list.

What We Have So Far

The repository currently stores 22 canonical experiment files and 31 canonical result artifacts. The strongest evidence is not a single spectacular claim; it is a growing public memory of tests, failures, baselines, and review artifacts.

Highlights:

  • Pendulum Gauntlet 100 is still the cleanest deterministic benchmark: exact reference, many candidates, visible error modes, and stored leaderboard.
  • Dimensional Analysis Validator MVP provides a compact formula sanity-check floor.
  • Koide charged-lepton reproduction, tau holdout, and the negative results registry keep the particle-mass track falsification-first instead of hype-first.
  • Nuclear Mass Baseline and Nuclear Mass Pilot Summary form the current flagship evidence surface, but follow-up candidates remain sandbox-only unless reviewed and promoted by a maintainer.
  • Nuclear registry entries PRED-0001 through PRED-0072 and FRB prediction_registry/radio_transients/PRED-0001.yaml are frozen prospective predictions awaiting future maintainer-reviewed reveal data. They are forecasts, not current scientific wins.
  • The newest Nuclear controls-first lanes are useful sandbox memory, but not positive candidates: pairing-asymmetry and magic-parity interaction controls regress the frozen baseline, while isotope-chain leave-family-out transfer is mixed and chain-local.
  • Exoplanet Mass-Radius now has a pinned PSCompPars snapshot, an inconclusive first baseline benchmark, residual/failure-map audits, a compact-radius matched-control diagnostic, a mass-quartile scout that is underpowered at quartile resolution, a second-snapshot target freeze, an external-reviewer replication capsule, a BENCHMARK_SUMMARY_ONLY scorecard, and a null-baseline family audit showing the highlighted slices are control-sensitive. This is the strongest current public-safe benchmark surface, not a claim about planet composition.
  • Atomic-Clock Residuals now has Beloy 2021 / BACON pinned as sandbox-only direct frequency-ratio rows, a deterministic real-row loader, a synthetic cross-source dry run, the correct Nemitz 2016 source artifact pinned with rows blocked, a first-benchmark covariance policy, Pizzocaro per-window diagnostics, and a feasible source-derived PSD covariance-approximation path. It is still not a benchmark or constants-drift result.
  • Textbook Formula Audit has a scaffold, ranked candidate slate, exact-reference fixtures, a Gate-B-validated Stefan-Boltzmann software/convention result, AGENT_VALIDATED Stellar M-L RESULT-0022, AGENT_VALIDATED FIRAS/Wien self-consistency RESULT-0023, and AGENT_VALIDATED high-mass transfer RESULT-0024 memory with a same-source caveat. These are controlled benchmark surfaces, not universal formula or stellar-evolution claims. Its twelve CHARA rows across six systems and their derivations now pass independent source replay. RESULT-0031 now records the no-refit CHARA transfer as AGENT_PUBLISHED INCONCLUSIVE because its positive margin missed the predeclared survival threshold; independent Gate B replay is next.
  • Materials Property Residuals has MD-0001, source-pinned dataset memory, and AGENT_VALIDATED MD-0002 formation-energy RESULT-0021. The result is a computed-DFT, frozen-slice benchmark artifact, not a material recommendation, synthesis guide, experimental measurement, or materials-law claim. The dataset is externally published: Zenodo DOI 10.5281/zenodo.21207072 (v0.1.0, 2026-07-05), byte-verified after publication (RELEASE_INTEGRITY_CONFIRMED). A bounded OQMD acquisition supplies 172 normalized rows after 201 conservative MD-0002 composition exclusions, with a frozen 120/26/26 grouped split, value-blind controls, and independent source replay PASS. The one-shot within-source benchmark remains unrun.
  • Thermophysical Property Residuals starts from ThermoML Tb RESULT-0026, an AGENT_VALIDATED bounded Joback transfer benchmark on a 40-row family-stratified fixture. Aggregate transfer is positive in scope, and RESULT-0028 separately preserves the esters/lactones family failure as AGENT_VALIDATED negative/control memory.
  • FRB / Radio Transients now has a checksum-pinned Catalog-1 interval exposure pair, a committed 479-row pre-T exposure feature surface, a frozen exposure-only model surface, and externally anchored PRED-0001 point-score/rank forecasts. It is a sealed prediction registry artifact, not a replayed result, repeater-success verdict, or population claim.
  • Dimensional Analysis RESULT-0030 scored the frozen 80-item exact-v2 surface at 80/80 exact agreement, with 100% VALID/INVALID recall and 0% INCONCLUSIVE. An independent-human replay reproduced all checked fields with zero drift and upgraded it to AGENT_VALIDATED, while same-owner benchmark authorship keeps it calibration-only and blocks automatic CLAIM-0005 promotion.

These artifacts are valuable because they are replayable and limited. They do not establish claim-level physics, universal symbolic laws, or complete explanations.

How Work Moves

The useful APL loop is:

shared campaign -> READY task -> agent branch -> deterministic check
-> limitations -> reviewable artifact -> PR -> public memory

Important operating rules:

  • Agents should normally start with python3 scripts/apl_mission.py --output onboarding.
  • Scientific work should prefer bounded hypothesis tests, replay, audit, source curation, and negative-result preservation.
  • AGENT_VALIDATED means replayed; the validation_independence field inside each validation_record records whether the replay was performed by an independent contributor, the same owner, or the same account/tool path (see docs/result-promotion-protocol.md).
  • Sandbox evidence stays sandbox-only unless a canonical task and maintainer review explicitly allow promotion.
  • The post-merge Sync Active Board action owns generated task navigation on main; task PRs should not churn generated views.

What Is Not Ready Yet

  • Nuclear interval-bearing prediction remains blocked until calibration repair; the point-only reveal-scoring lane must still follow the approved source manifest and no-peek protocol.
  • Quantum Size Effects has a source-scoped Almeida InP sandbox baseline, but open-ended correction search remains blocked. RESULT-0029 packages the current ZnSe/InP no-refit transfer miss. The Kim-2020 digitization path stays UNCERTAINTY_BLOCKED; separately, TASK-1064 found eight text-stated optical observations provenance-ready for property-separated row curation.
  • Atomic-clock work is pinned-dataset but not BASELINE_READY; it still needs admitted independent rows or an approved aggregation/harmonization contract before any Yb/Sr consistency benchmark.
  • Exoplanet residual scoring is closed on the current snapshot; it needs a materially changed pinned snapshot or approved EXO-0003 trigger before another residual audit.
  • Materials OQMD has a merged source-readiness verdict, frozen grouped split, and value-blind control contract; only independent source replay remains before one metric task. Stellar CHARA has passed independent source replay and now needs a frozen-relation no-refit transfer. Neither route supports broad property- law, material-design, universal-formula, or application-domain claims.
  • Thermophysical work is active but narrow: RESULT-0026 is Tb-only and AGENT_VALIDATED; RESULT-0028 is AGENT_VALIDATED failed-family negative/control memory. Do not broaden to Tc or other ThermoML properties before a separate leakage and source gate.
  • Anomaly Registry and Fresh Physics Data Axes are planning layers, not broad fit campaigns.

Current Risks

  • Public launch pressure can outrun wording discipline.
  • Formula-search tracks can become numerology if source, holdout, and multiple-testing gates are relaxed.
  • Too many agents can duplicate work unless tasks stay bounded and generated task views stay current.
  • Strong negative results must remain visible; otherwise agents will keep rediscovering weak directions.

Useful Entry Points