Skip to content

[S1] Positioning + netscript-bench #302

Description

@rickylabs

Child of the Road-to-0.0.1-stable umbrella. Owns positioning measurement + netscript-bench.

Scope

Build netscript-bench as the measurement instrument that locks and defends NetScript's position (AI-agent build-efficiency). Fork Encore's ai-benchmark protocol: black-box vitest HTTP probes, composite score (test_pass_rate / production_rubric_pass_rate / turns_to_green / cost / lines_of_code), layered tasks from NetScript's tutorial spine (t1-storefront-api, t2-saga-queue-cron, t3-reactive-dashboard + t3b reactive Fresh-UI extended axis). Full design in validated netscript-bench-research.v2.md.

Sequencing

  1. NetScript self-bench first (needs no user decisions): harness + scorer + t1 + t2 as a regression detector on every release. This is a beta gate.
  2. Cross-framework batch + published leaderboard: BLOCKED ON user decisions D1-D7 (VM provider, framework list, task count, weights, model+cadence, repo location, license).

Acceptance criteria

  • Self-bench runs t1+t2 green with test_pass_rate >= 0.90 median; captures all artifacts (diff, transcript, vitest.json); composite scorer with editable weights.
  • (stable) Full core-5 batch published; NetScript composite top-2; reproduced across two model baselines; continuous per-release cadence.

Notes

Repo location is decision D6 (new rickylabs/netscript-bench vs subdir). Per-package benchmark port from netscript-start/benchmark/ is a separate later session. Corpus-overlap confound (agent training data includes NetScript docs) disclosed in methodology.


Parent epic: #301

Part of #301

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions