Child of the Road-to-0.0.1-stable umbrella. Owns positioning measurement + netscript-bench.
Scope
Build netscript-bench as the measurement instrument that locks and defends NetScript's position (AI-agent build-efficiency). Fork Encore's ai-benchmark protocol: black-box vitest HTTP probes, composite score (test_pass_rate / production_rubric_pass_rate / turns_to_green / cost / lines_of_code), layered tasks from NetScript's tutorial spine (t1-storefront-api, t2-saga-queue-cron, t3-reactive-dashboard + t3b reactive Fresh-UI extended axis). Full design in validated netscript-bench-research.v2.md.
Sequencing
- NetScript self-bench first (needs no user decisions): harness + scorer + t1 + t2 as a regression detector on every release. This is a beta gate.
- Cross-framework batch + published leaderboard: BLOCKED ON user decisions D1-D7 (VM provider, framework list, task count, weights, model+cadence, repo location, license).
Acceptance criteria
- Self-bench runs t1+t2 green with
test_pass_rate >= 0.90 median; captures all artifacts (diff, transcript, vitest.json); composite scorer with editable weights.
- (stable) Full core-5 batch published; NetScript composite top-2; reproduced across two model baselines; continuous per-release cadence.
Notes
Repo location is decision D6 (new rickylabs/netscript-bench vs subdir). Per-package benchmark port from netscript-start/benchmark/ is a separate later session. Corpus-overlap confound (agent training data includes NetScript docs) disclosed in methodology.
Parent epic: #301
Part of #301
Child of the Road-to-0.0.1-stable umbrella. Owns positioning measurement +
netscript-bench.Scope
Build
netscript-benchas the measurement instrument that locks and defends NetScript's position (AI-agent build-efficiency). Fork Encore's ai-benchmark protocol: black-box vitest HTTP probes, composite score (test_pass_rate/production_rubric_pass_rate/turns_to_green/cost/lines_of_code), layered tasks from NetScript's tutorial spine (t1-storefront-api, t2-saga-queue-cron, t3-reactive-dashboard + t3b reactive Fresh-UI extended axis). Full design in validatednetscript-bench-research.v2.md.Sequencing
Acceptance criteria
test_pass_rate>= 0.90 median; captures all artifacts (diff, transcript, vitest.json); composite scorer with editable weights.Notes
Repo location is decision D6 (new
rickylabs/netscript-benchvs subdir). Per-package benchmark port fromnetscript-start/benchmark/is a separate later session. Corpus-overlap confound (agent training data includes NetScript docs) disclosed in methodology.Parent epic: #301
Part of #301