Skip to content

Research feedback: validate LakatoTree against preregistration, falsification, and autonomous-science baselines #10

Description

@gj3447

Context

This is external research/engineering feedback after reviewing the public repository at ac3bca4.

LakatoTree's clearest potential contribution is not any one ingredient, but the executable combination of:

  • a prediction and metric locked before measurement;
  • an external-measurement contract;
  • a pure deterministic judge;
  • a portable evidence record that forbids the producer from authoring its own verdict;
  • branching research-programme history with progressive/partial/equivalent/rejected outcomes;
  • replayable receipts plus a machine-checked kernel model.

Adjacent systems cover important subsets:

I did not find one public system that clearly combines all of the LakatoTree constraints above. However, absence of a close product match is not yet evidence that the combined method improves research reliability.

Main concern

The current honest boundary says that the default live path seals a client-submitted numeric value. This gives reproduction-confirmation, not measurement ownership. A deterministic verdict can therefore be perfectly replayable while still being derived from a biased, stale, or fabricated measurement.

Also, the Lakatos verdict layer is currently a designed governance rule. Its benefit over simpler preregistration plus an independent scorer has not yet been demonstrated experimentally.

Suggested decisive benchmark

Build a blinded benchmark in which multiple research agents receive the same objectives, data access, budgets, and injected failure cases.

Suggested arms:

  1. Ungoverned autonomous research agent.
  2. OSF-style preregistration record plus ordinary experiment tracking.
  3. POPPER-style sequential falsification.
  4. LakatoTree with client-submitted measurement.
  5. LakatoTree with signed independent producer or server-replayed measurement.

Suggested outcomes:

  • false-progressive rate on true nulls and poisoned measurements;
  • acceptance rate of post-hoc metric/threshold changes;
  • detection rate for self-authored or forged verdicts;
  • independent replay and receipt-rederivation success;
  • calibration of progressive/partial/equivalent/rejected outcomes;
  • time, compute, and authoring overhead;
  • robustness when the same target is repeatedly confirmed or when branches share evidence.

Pre-register the benchmark itself outside the evaluated LakatoTree instance, freeze hidden task seeds, and publish all failed and abandoned branches.

Possible deliverables

  • docs/RELATED_WORK.md with a feature-by-feature claim table;
  • a small public benchmark corpus containing null, positive, poisoned, and ambiguous cases;
  • an explicit measurement-assurance ladder: client assertion → signed producer → independent replay → independently observed world state;
  • a paper-level claim no broader than the observed reduction in false progress or post-hoc rationalization.

Why this matters

If LakatoTree beats the simpler baselines, the contribution becomes much sharper: not merely "Lakatos implemented in software," but an empirically validated governance kernel that reduces false scientific progress in autonomous research workflows.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions