Program
Independent reproduction for the 2026-09-04 RuV SOTA cycle.
Tracks:
Independent roles
- Research freezes source claims, versions, licenses, datasets, hardware assumptions, and exclusions.
- Baseline builds current unmodified RuV behavior and immutable evaluator artifacts.
- Implementation builds only adapters needed to exercise candidate primitives.
- Adversarial review designs mutation, replay, substitution, degradation, and distribution-shift cases without editing the candidate.
- Security validates authority boundaries, tenant isolation, privacy, supply-chain assumptions, and rollback.
- Testing runs deterministic unit, integration, resource-exhaustion, malformed-input, and partial-failure cases.
- Reproducibility pins model, harness, dataset, seeds, environment, costs, and exact commands.
- Release recommends ACCEPT, REJECT, or INCONCLUSIVE without merge authority.
Track A: HookPry
Compare current plugin or hook admission behavior with the RVM hook-update gate. Use ephemeral synthetic plugins only. Freeze benign and malicious update sequences before results. Measure unauthorized hook bindings, rights widening, legitimate acceptance, validation p50/p95, CPU, receipt size, and failure modes.
Track B: PatchBench
Compare crash-only, crash-plus-regression, and root-cause validation on seeded vulnerability repairs. Include transformed attack cases and patch-similarity analysis. Measure apparent versus true solve rate, false promotion, utility regressions, evaluator time, and model cost.
Track C: UMPeek
Compare no personalization, current RuV personalization, fixed probing, adaptive probing, response defense, and stateful counterfactual defense where feasible. Measure recovery F1, false inference, personalization utility, latency, tokens, model cost, and cross-tenant leakage.
Governance
Evaluators and attack sets freeze before candidate outcomes are visible. Candidate branches cannot alter evaluator, RVM authority, credentials, or production data. No real third-party plugin ecosystem attacks, credential extraction, deployment, autonomous merge, or irreversible migration.
Acceptance
Each track must publish baseline, candidate, workload, versions, seeds, sample size, absolute and relative results, variance, ablations, regressions, overhead, cost, exact reproduction steps, and negative results. Originating-team claims remain unverified until matched RuV reproduction completes.
Program
Independent reproduction for the 2026-09-04 RuV SOTA cycle.
Tracks:
Independent roles
Track A: HookPry
Compare current plugin or hook admission behavior with the RVM hook-update gate. Use ephemeral synthetic plugins only. Freeze benign and malicious update sequences before results. Measure unauthorized hook bindings, rights widening, legitimate acceptance, validation p50/p95, CPU, receipt size, and failure modes.
Track B: PatchBench
Compare crash-only, crash-plus-regression, and root-cause validation on seeded vulnerability repairs. Include transformed attack cases and patch-similarity analysis. Measure apparent versus true solve rate, false promotion, utility regressions, evaluator time, and model cost.
Track C: UMPeek
Compare no personalization, current RuV personalization, fixed probing, adaptive probing, response defense, and stateful counterfactual defense where feasible. Measure recovery F1, false inference, personalization utility, latency, tokens, model cost, and cross-tenant leakage.
Governance
Evaluators and attack sets freeze before candidate outcomes are visible. Candidate branches cannot alter evaluator, RVM authority, credentials, or production data. No real third-party plugin ecosystem attacks, credential extraction, deployment, autonomous merge, or irreversible migration.
Acceptance
Each track must publish baseline, candidate, workload, versions, seeds, sample size, absolute and relative results, variance, ablations, regressions, overhead, cost, exact reproduction steps, and negative results. Originating-team claims remain unverified until matched RuV reproduction completes.