Causal inference, pre-registered experimentation, and adversarial verification for autonomous software engineering loops.
- What Is This Project?
- Why Does This Exist?
- Research Questions
- System Architecture
- Visual Architecture Diagrams
- Core Modules — What Each File Does
- Protocol Documents (16 Specifications)
- Test Suite & Benchmarks
- Verified Experimental Findings & Negative Results
- Design Decisions (6 Canonical Decisions)
- Protocol Deviations (10 Documented)
- Data & Artifacts
- What Was Done (Development Timeline)
- What Was NOT Done (Planned Work)
- Reproducibility & Quick Start
- Project Structure
- Research Roadmap
- License & Citation
LERM (Loop Engineering Research Machine) is an open-source scientific research machine that tests whether autonomous coding agent "loops" (verify-repair cycles) actually work — or whether published improvements are statistical artifacts of confounded experiments.
It treats the agent scaffold as an experimental intervention within a randomized controlled trial (RCT) framework, with:
- Cryptographic pre-registration (hypotheses sealed before execution)
- Fail-closed data firewalls (historical task contamination blocked mechanically)
- Adversarial skeptic auditor (5-attack vector suite from an independent model family)
- Compute-matched comparison arms (identical turn, token, and wallclock budgets)
- Unbiased pass^k reliability metric (not the inflated pass@k used in most papers)
The entire system runs at ₹0 / $0 cloud cost using local Ollama inference and free-tier OpenRouter APIs.
Autonomous software engineering agents are frequently claimed to be "dramatically improved" by in-loop verification, multi-agent debate, or automated retry. However, the majority of published results suffer from severe methodological confounds:
| Confound | What Happens | How LERM Prevents It |
|---|---|---|
| Compute Asymmetry | Treatment arm gets 2×–5× more tokens/time than control | Strict matched budgets (3 turns, 16K tokens, 120s wallclock) |
| Benchmark Leakage | Public benchmark tasks are in model pre-training data | Procedural AST mutations on fresh repositories (SWE-smith) |
| Stochastic Cherry-Picking | Authors report pass@k (any 1 of k worked) as success | Unbiased pass^k estimator (ALL k must succeed) |
| Ungrounded Verifiers | Verifiers game test fixtures, declare false accepts | Independent container-isolated evaluator; agent never sees tests |
| Post-Hoc Task Selection | Tasks pruned after inspecting model performance | Pre-registered task manifests; SHA-256 sealed before execution |
| ID | Research Question | Status |
|---|---|---|
| RQ1 | Does in-loop verification and repair (loop_verify) yield a statistically significant improvement in pass^k verified success over a compute-matched naive retry baseline (loop_retry) under identical turn, token, and wallclock caps? |
OPEN — awaiting EXP-LOOP-003 primary trial |
| RQ2 | How does the verifier False Accept Rate (FAR) degrade multi-turn repair trajectories on long-horizon software engineering tasks? | OPEN |
| RQ3 | Can an adversarial auditor (L0) from an independent model family refute spurious empirical findings before they enter the research knowledge base? | VERIFIED — L0 skeptic refuted 3 of 3 flawed findings |
LERM is structured into four distinct planes separated by strict causal firewalls:
┌─────────────────────────────────────────────────────────────────────────────┐
│ LAYER 4: SCIENTIFIC GOVERNANCE & ADVERSARIAL SKEPTIC (L0) │
│ • Skeptic 5-Attack Suite • Knowledge Base (KB) • Spend & Drift Guardrails │
└──────────────────────────────────────┬──────────────────────────────────────┘
│ Pre-reg Verification & Skeptic Review
┌──────────────────────────────────────▼──────────────────────────────────────┐
│ LAYER 3: CAUSAL CONTROL PLANE & DATA FIREWALL │
│ • Pre-registration Hasher • Fail-Closed Firewall • Deterministic Controller │
└──────────────────────────────────────┬──────────────────────────────────────┘
│ Matched Arms Dispatch & Isolation
┌──────────────────────────────────────▼──────────────────────────────────────┐
│ LAYER 2: L1 TRACE TELEMETRY & CONFOUND BALANCER │
│ • Unmerged 3-Outcome Traces • 10-Confound Checklist • Pass^k Estimator │
└──────────────────────────────────────┬──────────────────────────────────────┘
│ Ephemeral Container Mounts & Signals
┌──────────────────────────────────────▼──────────────────────────────────────┐
│ LAYER 1: EXECUTION PLANE & CONTAINER SANDBOX (ADR-0001) │
│ • Ephemeral Docker Sandbox • Read-Only Pytest Evaluator • Resource Manager │
└─────────────────────────────────────────────────────────────────────────────┘
Governing Principle:
TRUTH > CAUSAL VALIDITY > INDEPENDENCE > STATISTICAL POWER > REPRODUCIBILITY > THROUGHPUT
For the complete technical specification, see docs/ARCHITECTURE.md.
All diagrams are located in docs/diagrams/ as production SVG files:
| Diagram | Description | File |
|---|---|---|
| System Architecture | 4-layer separated plane overview | 01_system_architecture_3d.svg |
| Data & Control Flow | End-to-end experiment lifecycle from pre-reg to KB settlement | 02_data_control_flow_3d.svg |
| Agent Loop Comparison | Matched-budget loop_retry vs loop_verify (3 turns, 16K tokens, 120s) |
03_agent_loop_3d.svg |
| Verification Chain | 8-stage candidate verification: Baseline → Mutation → GT Fail → Repair → GT Pass → Reset Hash | 04_verification_eval_3d.svg |
| Failure & Recovery | Non-merged 3-outcome taxonomy (verified_success, infra_failure, model_failure) with OOM eviction | 05_failure_recovery_flow_3d.svg |
| Roadmap & Vision | 5-phase progression from control plane to autonomous AI scientist | 06_roadmap_vision_3d.svg |
| Module | Lines | Purpose |
|---|---|---|
stats.py |
301 | Statistical core. Unbiased pass^k (combinatorial C(c,k)/C(n,k)) and pass@k (Chen et al.) estimators, paired bootstrap CI (resamples tasks not runs), O'Brien-Fleming sequential stopping boundaries, power analysis, Wilson score & Clopper-Pearson intervals, calibration admission gates. |
skeptic.py |
218 | Adversarial Skeptic (L0). 5-attack suite: (1) replication with new seeds, (2) confound search, (3) shortcut/reward-hack hunt, (4) alternative explanation, (5) noise check against pre-registered CI threshold. Hard-enforces different model family at construction. |
firewall.py |
230 | Fail-closed data firewall. Quarantines historical holdouts (HOLDOUT-001–004) by task ID AND fixture content SHA-256 hash. Blocks string booleans, missing flags, metadata tampering, and copied fixture artifacts. |
candidate_pool.py |
266 | Candidate pool manager & rejection taxonomy. 9-gate admission pipeline: firewall → schema → provenance → treatment purity → hash verification → mutation dedup → diff dedup → config dedup → 8-step chain verification. 16 rejection categories. |
candidate_schema.py |
155 | Immutable candidate task unit schema. 20 typed fields (task_id format, git SHA validation, mutation rule allowlist, SHA-256 hash format, provenance requirements). Frozen at Schema v1.0.0. |
calibration.py |
155 | Phase D difficulty calibration. Two-stage sequential screening (N₁=10 Stage 1, N=20 Stage 2), guardband admission rule (0.35 ≤ p̂ ≤ 0.65), Wilson CI intersection with [0.30, 0.70], manifest YAML generator. |
confounds.py |
~135 | 10-dimension mechanical confound checklist. Validates: equal wallclock, equal tokens, equal attempts, equal cost, same model version, no model drift, comparable task order, same seed policy, same task set, no human intervention. |
trace.py |
~225 | Immutable trace store. Append-only JSONL. Non-merged 3-field outcome taxonomy: verified_success, infra_failure, model_failure. Tracks LERM framework overhead ratio (≤10%). |
prereg.py |
~145 | Cryptographic pre-registration. SHA-256 sealed YAML manifests. prereg.verify() asserts git-tracked files unmodified post-hoc. |
kb.py |
~215 | Knowledge Base & settlement. Three-valued verdicts: SUPPORTED, REFUTED, UNCERTAIN. Rejects hedging prose ("promising", "suggests", "trend"). Requires explicit unexplained field. |
controller.py |
~95 | Deterministic loop controller. Manages loop_retry (Arm A) vs loop_verify (Arm B) under matched budgets: max 3 turns, 16K tokens, 120s wallclock. |
evaluator.py |
~95 | Independent evaluator. Runs pytest in ephemeral container. Agent's self-reported claim is NEVER ingested as ground truth. |
resource.py |
~100 | Hardware resource manager. Monitors physical RAM via Windows API. Proactively evicts Ollama models (keep_alive=0) when RAM < 2GB to prevent OOM thrashing on 16GB hardware. |
guardrails.py |
~140 | Spend & drift guardrails. Three-tier spend ceiling ($25/hr, $200/day, $800/week). Kill-switch via state/KILL file. Canary prompt monitoring for silent API-side model weight updates. |
digest.py |
~95 | Daily/weekly digest generator. Summarizes trace telemetry into human-readable markdown reports. |
ledger.py |
~65 | Append-only experiment ledger. JSONL-backed immutable record of all experiment events. |
regression.py |
~95 | Regression detection. Compares current results against prior baselines to detect performance degradation. |
repairability.py |
~130 | Treatment-independent repairability classifier. Categories: REPAIRABLE, UNREPAIRABLE, INFRASTRUCTURE, AMBIGUOUS, EVALUATOR_FAILURE. Zero reliance on loop outcomes. |
trace_purity.py |
~50 | Calibration trace purity auditor. Verifies zero loop feedback / verifier diagnostics leaked into calibration traces. |
checkpoint.py |
~95 | Experiment checkpoint manager. Persists and restores in-flight experiment state for crash recovery. |
| Module | Lines | Purpose |
|---|---|---|
openhands.py |
266 | OpenHands agent-server adapter. HTTP client (stdlib urllib only) for OpenHands v1 & legacy v0 APIs. Conversation lifecycle: start → poll → events → sandbox. OmniRoute model plane integration with traced completions. Rate-limited retry with exponential backoff. |
All protocols are in docs/protocol/ and cryptographically hashed in the Phase C Audit Report:
| # | Protocol | File | What It Defines |
|---|---|---|---|
| 01 | Candidate Schema | 01_candidate_schema.md |
20-field immutable task unit: types, formats, validation rules |
| 02 | Repository Sampling | 02_repository_sampling_policy.md |
Deterministic repo selection (PRNG seeded), baseline pass requirement |
| 03 | Mutation Policy | 03_mutation_policy.md |
5 allowed AST mutation operators, prohibited operators, compound mutations |
| 04 | Generator Config | 04_candidate_generation_config.md |
SWE-smith invocation parameters, deterministic RNG seeding |
| 05 | Random Seed Policy | 05_random_seed_policy.md |
Master seed hierarchy (MASTER_SEED=20260921), derived subsystem seeds |
| 06 | Calibration Protocol | 06_calibration_protocol.md |
Two-stage sequential screening (N₁=10, N=20), guardband admission |
| 07 | Repairability Protocol | 07_repairability_protocol.md |
Treatment-independent classification, reference patch verification |
| 08 | Ground Truth Protocol | 08_ground_truth_protocol.md |
Container-isolated pytest execution, zero agent self-report trust |
| 09 | Rejection Taxonomy | 09_candidate_rejection_taxonomy.md |
16 deterministic rejection categories with mandatory logging |
| 10 | Manifest Format | 10_candidate_manifest_format.md |
Frozen YAML manifest structure, atomic validation |
| 11 | Provenance Specification | 11_provenance_specification.md |
Required provenance keys: generator_version, config_hash, timestamp |
| 12 | Independence Specification | 12_independence_specification.md |
One task per function, triple dedup (mutation_id, diff hash, config tuple) |
| 13 | Leakage Audit | 13_information_leakage_audit.md |
Forbidden treatment keys, zero generator→trace dependencies |
| 14 | Adversarial Test Suite | 14_adversarial_test_suite.md |
20 attack vectors specified for candidate pool integrity |
| 15 | Reproducibility Tests | 15_reproducibility_test_results.md |
Bit-for-bit determinism across 3 configurations |
| 16 | Calibration OC Analysis | 16_calibration_operating_characteristics.md |
Exact binomial admission probabilities across true-p values |
Total: 88 tests, all passing (verified 2026-10-01, Python 3.14, pytest 9.1.1, 4.79s)
| File | Tests | What It Validates |
|---|---|---|
tests/test_core.py |
27 | Core statistics (pass^k, pass@k, bootstrap CI, power analysis, Fisher exact p, O'Brien-Fleming boundaries), pre-registration seal/verify, Knowledge Base settlement (hedging rejection, unexplained field), verifier quality FAR/FRR, trace 3-outcome taxonomy, confound checklist, regression detection, controller budget enforcement, Wilson score interval, calibration sample size |
tests/test_firewall.py |
13 | Data firewall: historical task rejection (by ID and fixture hash), metadata tampering detection, string boolean rejection, missing flag rejection, copied fixture detection, fresh candidate admission, trace dataset contamination blocking |
tests/test_candidate_adversarial.py |
20 | Full adversarial audit — 20 distinct attack vectors against the candidate pool (see §8.1 below) |
tests/test_repairability.py |
7 | Repairability classifier: repairable, unrepairable (no patch, no signal), infrastructure failure, evaluator syntax failure, ambiguous/flaky, timeout |
tests/test_skeptic_planted.py |
7 | Planted-finding skeptic attacks: confounded finding killed, shortcut finding killed, noise finding killed, honest finding survives, same-model-family refused, seed policy flagged, fixture reset mismatch flagged |
tests/test_calibration_stage2.py |
14 | Stage 2 precision evaluation: admission gate (7 ≤ X ≤ 13), Wilson CI intersection, edge cases (0/20, 20/20, 7/20, 13/20), metrics computation, trace purity |
Every vector validates fail-closed rejection of invalid, poisoned, or contaminated candidates:
# Attack Vector Expected Rejection
1 Historical task injected into calibration HISTORICAL_FIREWALL_BREACH
2 Historical task injected into primary HISTORICAL_FIREWALL_BREACH
3 Fresh task with fake provenance PROVENANCE_TAMPERED
4 Fresh task with missing provenance CandidateValidationError
5 Fresh task with altered provenance PROVENANCE_TAMPERED
6 Candidate copied from historical fixture FirewallViolation
7 Candidate with one-byte source perturbation HASH_MISMATCH
8 Duplicate mutation DUPLICATE_MUTATION
9 Same mutation under different ID DUPLICATE_DIFF
10 Treatment-success leaked into metadata TREATMENT_LEAKAGE
11 Calibration result leaked into generator state TREATMENT_LEAKAGE
12 Candidate missing mutation ID CandidateValidationError
13 Candidate mismatched repository/commit CandidateValidationError
14 Candidate mismatched content hash CandidateValidationError
15 Candidate baseline already fails BASELINE_FAILURE
16 Candidate mutation produces no failure MUTATION_NO_EFFECT
17 Candidate failure is infrastructure-only INFRASTRUCTURE_FAILURE
18 Reference repair does not restore PASS REFERENCE_REPAIR_FAILURE
19 Candidate reset hash differs RESET_FAILURE
20 Candidate generated with same seed twice DUPLICATE_CONFIG
100% bit-for-bit deterministic SWE-smith bug generation verified across 3 independent configurations:
- CONFIG_A (
mewwts__addict, seed=42): 2 files identical - CONFIG_B (
mewwts__addict, seed=101): 6 files identical - CONFIG_C (
mewwts__addict, seed=999): 18 files identical
Run: python verify_3_configs_reproducibility.py
- Power Analysis: At N=7 tasks (k=5, p₀=0.50^5=0.03125, α=0.05), power to detect MDE=0.20 is 18.62%. N≥44 required for 80% power at MDE=0.20.
- Calibration Operating Characteristics: Guardband rule (0.35 ≤ p̂ ≤ 0.65) at true p=0.50 yields P(Admit)=0.8770; at true p=0.05 yields P(Admit)=0.0000 (zero floor leakage).
- Wilson Score vs Wald: Wilson score CI is preferred because Wald severely under-covers near boundaries.
LERM mandates public documentation of all results including negative ones. This is central to scientific integrity.
- File:
findings/LE-0001.yaml - Hypothesis: Independent verification in the loop improves pass^5 verified success more than equal-compute retries.
- Observed Effect: +0.175 [+0.067, +0.282], FAR=0.051
- Skeptic Verdict:
UNCERTAIN— survived confound search, shortcut hunt, and noise check; flaggedNEEDS_EXECUTIONfor replication and alternative explanation. - Model:
qwen/qwen3-coder-480b(simulated), N=1200 runs, 60 tasks
- File:
negative_results/FINDING-EXP-LOOP-001.yaml - Status:
REFUTED - Reason: Confound search KILLED (wallclock 33% imbalanced, tokens 9% imbalanced). Noise check KILLED (CI includes zero at n_tasks=4).
- Model:
inkling-small:free
- File:
negative_results/FINDING-EXP-LOOP-002.yaml - Status:
REFUTED - Finding: Zero treatment variance. Out of 79 valid trials, 100% of successful trials solved on Turn 1 before retry or verification mechanics could engage.
- Root Cause: Hand-crafted holdout tasks were too easy for the model, collapsing multi-turn variance.
- Remediation: Designed the Phase C Candidate Generation and Sequential Calibration Protocol to enforce 0.35 ≤ p̂ ≤ 0.65 single-shot baseline difficulty.
- File:
negative_results/NEG-LIVE-0001.yaml - Status:
REFUTED - Finding: 6 pipeline defects discovered: (D1) hardcoded demo data in regression_result, (D2) model emitted raw JSON without ActionEvent, (D3) HOLDOUT-001 passed on untouched fixtures, (D4) checkers missing exit code 2, (D5) trace timing 52μs apart, (D6) shortcut_pilot mock placeholder.
- Remediation: All 6 defects catalogued and fixed. Trust pipeline rebuilt from scratch.
All decisions are recorded in docs/DECISIONS.md:
| ID | Question | Resolution |
|---|---|---|
| DECISION-0001 | Governing protocol version | Adopted LERM Protocol v2 (SHA-256 sealed) |
| DECISION-0002 | EXP-LOOP-003 candidate pool scope | Phase D-partial on addict only (K_max=14). marshmallow and sqlfluff blocked by offline dependency failures. |
| DECISION-0003 | Execution mechanism parity | Raw completion calibration sets upper bound on intrinsic difficulty; EXP-LOOP-003 must match pathways. |
| DECISION-0004 | Raw completion vs OpenHands calibration | Raw completion as cheap screen; OpenHands confirmation gate mandatory before primary admission. qwen2.5-coder:1.5b cannot execute OpenHands tool actions (0/20 trials). |
| DECISION-0005 | Agent-under-test model | openrouter/thinkingmachines/inkling-small:free confirmed. qwen2.5-coder:1.5b/3b disqualified (cannot invoke bash/editor tools in OpenHands). |
| DECISION-0006 | Ceiling effect resolution | In-agent bash testing is valid capability. Candidates require compound/multi-mutation. Calibration throttled ≤15 trials/day for ₹0 budget preservation. |
All deviations are recorded in docs/DEVIATIONS.md with evidence:
n_reruns >= 2k, notn_reruns == k(at n=k=5, task-level CI blows up)- Bootstrap resamples tasks, not runs (tasks are the unit of generalisation)
- DuckDB is optional (JSONL is source of truth; SQLite fallback)
- Skeptic attacks 1 & 4 return
NEEDS_EXECUTION, neverPASS(pre-execution honesty) - Spend cap default deliberately low ($3/h, $40/day)
- Skeptic model family check is a hard error at construction
- PRNG_SEED_REPO reconciled to 61022 (from master seed hierarchy)
- Repository pool restricted to
addictonly (DECISION-0002) - Golden-copy preservation wrapper for SWE-smith
shutil.rmtreebehavior - Autonomous tool interaction & compound mutation calibration policy (DECISION-0006)
| File | Experiment |
|---|---|
preregistrations/LE-0001.yaml |
Offline pilot simulation |
preregistrations/EXP-LOOP-001.yaml |
Live pilot with inkling-small |
preregistrations/EXP-LOOP-002.yaml |
Live pilot with wallclock matching |
| File | Description |
|---|---|
holdout/tasks/HOLDOUT-001.yaml – HOLDOUT-004.yaml |
Historical holdout tasks from EXP-LOOP-002 (quarantined in firewall) |
holdout/tasks/_schema.yaml |
Holdout task schema definition |
| File | Description |
|---|---|
data/phase-d/candidates.jsonl |
7 generated candidate tasks (Cohort 1) |
data/phase-d/compound_candidates.jsonl |
Compound multi-mutation candidates |
data/phase-d/candidates_manifest.yaml |
Frozen candidate manifest |
data/phase-d/calibration_traces*.jsonl |
Calibration trace records |
data/phase-d/calibration_summary*.json |
Summary metrics for calibration runs |
data/phase-d/admitted_primary_manifest.yaml |
Tasks admitted to primary pool |
| File | Description |
|---|---|
reports/PHASE_C_AUDIT_REPORT.md |
Phase C Protocol Freeze — 20-gate binary checklist, 26 cryptographically hashed artifacts, all gates PASS |
reports/P0_LIVE_VERIFICATION.md |
P0 live pipeline verification report |
reports/GITHUB_PUBLICATION_REPORT.md |
Publication audit |
reports/phase-d/ |
Phase D preflight, remediation design, reconciliation, dependency audit, final report |
- Built the full 4-layer architecture from scratch
- Implemented cryptographic pre-registration with SHA-256 sealing
- Implemented fail-closed data firewall with historical task quarantine
- Built adversarial skeptic (L0) with 5-attack vector suite
- Implemented 10-dimension mechanical confound checklist
- Built unbiased pass^k estimator (combinatorial, not naive power)
- Built paired bootstrap CI (task-level, not run-level)
- Implemented O'Brien-Fleming sequential stopping boundaries
- Built hardware resource manager with Ollama OOM eviction
- Built Knowledge Base with three-valued verdicts and hedging rejection
- Built spend guardrails with kill-switch
- Ran 40 trials (EXP-LOOP-001) → REFUTED (confounded wallclock & tokens)
- Ran 79 trials (EXP-LOOP-002) → REFUTED (ceiling effect, 100% Turn 1 solve)
- Discovered pipeline integrity defects (NEG-LIVE-0001) → 6 defects catalogued
- Froze 16 protocol specifications
- Built candidate schema (20-field immutable CandidateTaskUnit)
- Built candidate pool manager with 9-gate admission pipeline
- Built 20-vector adversarial test suite → 20/20 PASS
- Verified SWE-smith bit-for-bit reproducibility across 3 configs
- Computed calibration operating characteristics (exact binomial)
- All 20 Phase C gate conditions PASS
- Cryptographically hashed all 26 protocol artifacts
- Generated Cohort 1 candidates (7 tasks) from
mewwts__addict - Discovered marshmallow/sqlfluff offline dependency failures
- Ran raw completion calibration → discovered ceiling effect with OpenHands tool access
- Designed compound mutation strategy (DECISION-0006)
- Generated compound candidate pairs
- Built Phase D calibration evaluation module
- Built OpenHands adapter (stdlib-only HTTP, v1+v0 API support)
- Probed model capabilities:
qwen2.5-coder:1.5bdisqualified (0 tool actions),inkling-small:freeconfirmed - Built Stage 2 calibration tests (14 tests)
| Script | Purpose |
|---|---|
scripts/preflight.py |
P0 environment preflight checks |
scripts/demo_offline.py |
Full offline demo: prereg → traces → skeptic → KB → digest |
scripts/validate_tasks.py |
Holdout task schema validation |
scripts/run_causal_experiment.py |
Primary causal experiment runner (ready for EXP-LOOP-003) |
scripts/run_live_lerm.py |
Live LERM pipeline runner |
scripts/generate_phase_d_candidates.py |
Phase D candidate generation |
scripts/generate_compound_candidates.py |
Compound multi-mutation candidate generation |
scripts/run_phase_d_calibration.py |
Phase D calibration (raw completion) |
scripts/run_phase_d_calibration_openhands.py |
Phase D calibration (OpenHands) |
scripts/verify_broken_fixtures.py |
Fixture integrity verification |
scripts/verify_phase_a.py |
Phase A verification |
scripts/audit_phase_b.py |
Phase B audit |
| Item | Status | Why |
|---|---|---|
| EXP-LOOP-003 Primary Causal Trial | PLANNED |
Awaiting Phase D candidate pool completion. Pre-registered; script ready (scripts/run_causal_experiment.py). |
| Phase D Full Pool (K=30 across 3 repos) | BLOCKED |
marshmallow and sqlfluff have unpinned floating dependencies that fail pytest offline. Currently K_max=14 from addict only. |
| Multi-Cloud Ray Execution | VISION |
Phase 4/5 roadmap item. Requires SkyPilot cluster provisioning. |
| Autonomous 24/7 Research Loop | VISION |
Self-directed hypothesis generation, auto-pre-registration, paper draft synthesis. |
| LaTeX Manuscript Synthesis | VISION |
Automated paper generation from settled KB records. |
| DVC Integration | DEFERRED |
Large trace files tracked via git-ignored directories; DVC or release assets planned for scale. |
- Python 3.9+ (tested on Python 3.14, Windows 11 / WSL2)
- Docker Desktop / Engine (required for containerized evaluation)
- Git
- (Optional) Ollama for local zero-cost inference
# Clone the repository
git clone https://github.com/Nia00-glitch/lerm.git
cd lerm
# Install package and core dependencies in editable mode
pip install -e .
# (Optional) Install dev dependencies for analytics
pip install -e ".[dev]"# Copy example environment configuration
cp .env.example .env
# Edit .env with your local or remote LLM endpoint
# Works with local Ollama (http://127.0.0.1:11434) at zero monetary cost
# Or use OpenRouter free-tier models# 1. Run complete test suite (88 tests)
pytest tests/ -v
# 2. Validate holdout tasks against schema
python scripts/validate_tasks.py
# 3. Verify SWE-smith bit-for-bit determinism across 3 configurations
python verify_3_configs_reproducibility.py
# 4. Verify pre-registration hash integrity
python -c "from lerm.prereg import verify; print(verify('preregistrations/EXP-LOOP-002.yaml'))"
# 5. Run full offline demo (prereg → traces → skeptic → KB → digest)
python scripts/demo_offline.py
# 6. Run P0 preflight environment checks
python scripts/preflight.pymake preflight # P0 gate — environment health check
make selftest # P5 acceptance: skeptic must kill planted findings
make demo # Full offline end-to-end demo
make lock # Lock pip dependencies
make clean # Clean runtime artifactslerm/
├── lerm/ # Core Python package (21 modules)
│ ├── __init__.py
│ ├── stats.py # Statistical core (pass^k, CI, power, sequential testing)
│ ├── skeptic.py # Adversarial L0 skeptic (5-attack suite)
│ ├── firewall.py # Fail-closed data firewall
│ ├── candidate_pool.py # 9-gate candidate admission pipeline
│ ├── candidate_schema.py # 20-field immutable task unit schema
│ ├── calibration.py # Phase D two-stage calibration
│ ├── confounds.py # 10-dimension confound checklist
│ ├── trace.py # Immutable JSONL trace store
│ ├── prereg.py # Cryptographic pre-registration
│ ├── kb.py # Knowledge Base & settlement
│ ├── controller.py # Deterministic loop controller
│ ├── evaluator.py # Container-isolated pytest evaluator
│ ├── resource.py # Hardware RAM monitor & OOM eviction
│ ├── guardrails.py # Spend ceilings & kill-switch
│ ├── digest.py # Daily/weekly digest generator
│ ├── ledger.py # Append-only experiment ledger
│ ├── regression.py # Regression detection
│ ├── repairability.py # Treatment-independent classifier
│ ├── trace_purity.py # Calibration trace purity auditor
│ ├── checkpoint.py # Crash recovery checkpoint manager
│ └── adapters/
│ ├── __init__.py
│ └── openhands.py # OpenHands agent-server HTTP adapter
│
├── tests/ # Test suite (88 tests)
│ ├── test_core.py # 27 core tests
│ ├── test_firewall.py # 13 firewall tests
│ ├── test_candidate_adversarial.py # 20 adversarial attack tests
│ ├── test_repairability.py # 7 repairability tests
│ ├── test_skeptic_planted.py # 7 planted skeptic tests
│ └── test_calibration_stage2.py # 14 calibration tests
│
├── docs/
│ ├── ARCHITECTURE.md # Full technical architecture specification
│ ├── DECISIONS.md # 6 canonical design decisions
│ ├── DEVIATIONS.md # 10 documented protocol deviations
│ ├── SPEC.md # High-level specification
│ ├── ADR-0001-execution-plane.md # Architecture Decision Record
│ ├── protocol/ # 16 frozen protocol specifications
│ │ ├── 01_candidate_schema.md
│ │ ├── 02_repository_sampling_policy.md
│ │ ├── ...
│ │ └── 16_calibration_operating_characteristics.md
│ └── diagrams/ # 6 SVG architecture diagrams
│ ├── 01_system_architecture_3d.svg
│ ├── ...
│ └── 06_roadmap_vision_3d.svg
│
├── scripts/ # 12 executable scripts
│ ├── preflight.py # Environment health check
│ ├── demo_offline.py # Full offline demo
│ ├── run_causal_experiment.py # Primary causal experiment runner
│ ├── generate_phase_d_candidates.py
│ ├── generate_compound_candidates.py
│ ├── run_phase_d_calibration.py
│ ├── run_phase_d_calibration_openhands.py
│ └── ...
│
├── preregistrations/ # SHA-256 sealed experiment manifests
│ ├── LE-0001.yaml
│ ├── EXP-LOOP-001.yaml
│ └── EXP-LOOP-002.yaml
│
├── findings/ # Settled research findings
│ └── LE-0001.yaml # Offline pilot (UNCERTAIN)
│
├── negative_results/ # Mandatory negative result documentation
│ ├── FINDING-EXP-LOOP-001.yaml # REFUTED (confounded)
│ ├── FINDING-EXP-LOOP-002.yaml # REFUTED (ceiling effect)
│ └── NEG-LIVE-0001.yaml # REFUTED (pipeline defects)
│
├── holdout/ # Historical holdout tasks (quarantined)
│ └── tasks/
│ ├── HOLDOUT-001.yaml ... HOLDOUT-004.yaml
│ └── _schema.yaml
│
├── data/phase-d/ # Phase D calibration data & manifests
├── reports/ # Audit reports & verification evidence
│ ├── PHASE_C_AUDIT_REPORT.md # 20-gate Phase C freeze (all PASS)
│ ├── P0_LIVE_VERIFICATION.md
│ ├── GITHUB_PUBLICATION_REPORT.md
│ └── phase-d/ # Phase D reports
│
├── pyproject.toml # Package configuration (v1.0.0)
├── requirements.txt # Minimum dependencies
├── Makefile # Build targets
├── .env.example # Environment template
├── .gitignore # Comprehensive ignore rules
└── LICENSE # MIT License
| Phase | Status | Description |
|---|---|---|
| Phase 0/1 | COMPLETED |
Control plane foundation: pre-registration, firewall, skeptic, confounds, pass^k, KB, guardrails |
| Phase B | COMPLETED |
Pilot trials (EXP-LOOP-001, 002) — refuted, ceiling effect discovered |
| Phase C | COMPLETED |
Protocol freeze, adversarial audit (20/20), reproducibility verification, 26 artifacts hashed |
| Phase D | IN PROGRESS |
Candidate generation & sequential calibration on mewwts__addict |
| Phase 2/3 | PLANNED |
EXP-LOOP-003 primary causal trial (N≥20 tasks × k=5 reruns) |
| Phase 4/5 | VISION |
24/7 autonomous AI scientist: multi-cloud Ray execution, auto-pre-registration, LaTeX manuscript synthesis |
Distributed under the MIT License.
@software{lerm2026,
author = {Nia00-glitch},
title = {LERM: Loop Engineering Research Machine — Causal Inference and Adversarial Verification for Autonomous Agent Loops},
year = {2026},
url = {https://github.com/Nia00-glitch/lerm}
}