Skip to content

release: beta.10 wave — complete CLI coverage, agentic combo union, stabilization (integration → main)#800

Merged
rickylabs merged 20 commits into
mainfrom
feat/beta10-integration
Jul 17, 2026
Merged

release: beta.10 wave — complete CLI coverage, agentic combo union, stabilization (integration → main)#800
rickylabs merged 20 commits into
mainfrom
feat/beta10-integration

Conversation

@rickylabs

Copy link
Copy Markdown
Owner

Summary

The 0.0.1-beta.10 wave: complete CLI coverage + stabilization, the agentic-combo release union, honest CI for stacked waves, and the docs-site refresh — 17 evaluated PRs merged into feat/beta10-integration, every one with a supervisor-dispatched opposite-family IMPL-EVAL PASS and green CI.

This PR's merge auto-closes the milestone's remaining issues (all fixes verified on this branch):

Closes #762
Closes #763
Closes #773
Closes #774
Closes #781
Closes #782
Closes #783
Closes #785
Closes #791
Closes #792
Closes #796

(#769 and epic #721 already closed with evidence; #775/#778 re-milestoned to beta.13 with the paused dashboard.)

What ships (by sub-PR)

Area PRs
Specifier pinning + guard (p0) #770 (plugin CLI verbs, closes #763), guard proven-to-fail; repo-wide CI guard via the #715 union
Workers runtime #786 (health-check entrypoint + Flow-B fixture, closes #785), #793 (opt-in queue triggers, closes #792)
Fresh / fresh-ui #788 (render_ui recursion + generated-asset freshness CI gate, closes #773), #789 (Preact Windows dedupe, closes #782), #790 (markdown via Preact JSX runtime, −12% bundle, closes #783), #797 (clean-cache Signals resolution — CI-deterministic hydration builds)
Aspire/CLI generators #795 (7 emission fixes: executable capabilities, task flags, Vite keys, DB provider projection, bounded Garnet restore, SQLite workdir paths, abort flag — closes #791)
CI honesty #787 (integration-branch PRs run real lanes + lane visibility, closes #774), #772 (ts-suppression sweep 36→0, repo-drift gate CI-blocking, close-gate retry/fallback hardening — closes #762)
Harness/routing #794 (review-pairing ladder as data + Sol effort-selection doctrine), #776 (evaluator lane as data, closed models rejected in code), #777 (evaluator transport doctrine)
Release union #799 (main reconciled in: agentic combo #715, Fable restoration #784, OpenCode lane #779 — routing conflicts resolved per ratified doctrine, zero retired conditions)
Docs #771 (JSR taglines + byte-cap gate), #798 (docs-site refresh: full public command reference, agent/MCP/skills docs, tutorial revamp, stale sweep — closes #796)

Release readiness

After merge (owner runbook)

  1. Milestone 12 auto-closes its issues via the keywords above → close the milestone.
  2. Cut the release per netscript-release (release:cut → publish via GitHub Release/OIDC → race-free e2e-cli-prod verification). Not executed tonight by design.

Harness

Run dir: .llm/runs/beta10-cli--orchestrator/ (worklog, drift, per-slice verdicts, context-pack). Orchestrator: Claude · Fable 5 · low, session session_017LHrkXyMzsQwb9bqr82EFK. Merge of this PR is reserved for the owner.

🤖 Generated with Claude Code

https://claude.ai/code/session_017LHrkXyMzsQwb9bqr82EFK

rickylabs and others added 20 commits July 13, 2026 00:26
Point supervisor.md at the integration worktree cut from origin/main (0341c43).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…prompts

The entire night's record was uncommitted. Committing it, because the lesson this
run kept teaching is that an artifact on a disk is not an artifact that exists.

- MORNING-HANDOFF.md — the p0 (#769), the four owner decisions, and the pattern
  behind every defect found: we shipped things never checked against the thing
  they claim to control (#769 configs vs JSR · #773 embed vs source · #774 gates
  vs the PRs introducing them · NF1 policy vs the CLI it governs).
- worklog.md / drift.md — seven false-greens, each recorded with what it proved.
  Including the three claims of mine that were wrong (window.NSOne, the blanket
  no-reasoning caveat, "NF1 is on the PR") and how each was caught.
- canvas-prompts/P1..P6 — paste-ready, carrying the SVG-hole rule, the class-based
  contract, the real shipped CLI verbs, and the completion-report protocol.
- SCREEN-SPEC / PROPOSED-COMPONENTS / OPEN-QUESTIONS — the dashboard design contract.

Nothing merged, published, released, or closed. main untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
* chore(harness): establish #785 implementation evidence

* fix(workers): resolve rooted local job entrypoints once

* docs(harness): record #785 gate attribution

* test(scaffold): prove generic CLI jobs keep health pristine

* test(scaffold): prove Flow-B telemetry without health coupling

* docs(harness): record clean 60-gate acceptance
* chore(harness): prove the #774 CI repair plan

* chore(harness): clear the #774 plan gate

* fix(ci): run real lanes on integration pull requests

* chore(harness): record the #774 implementation verdict
* docs(harness): lock issue 773 implementation plan

* fix(fresh-ui): prove shipped render_ui recursion is bounded

* docs(harness): record issue 773 slice handoff
* chore(harness): lock #782 preact dedupe plan

* fix(fresh): canonicalize Preact module identity on Windows
* chore(harness): lock issue 783 markdown render plan

* fix(fresh-ui): render Markdown through Preact runtime

* docs(harness): record issue 783 gate evidence
* chore(harness): plan workers sample-trigger fix

* fix(workers): require explicit queue triggers

* test(workers): prove opt-in trigger scaffold behavior
…tion effort (#794)

Codify the owner-ratified (2026-07-16) adversarial review ladder into the
routing policy. Routing is data: CANONICAL_ROUTE_POLICY + the rendered
lane-policy.md view are updated together, referencing MODEL_IDS constants only.

Ladder (review of Codex/OpenAI-authored work, effort-paired):
- NEW light_implementation (Sol · low) -> review_codex_light: Opus 4.8 · high
  (token-limit fallback Sonnet 5 · high, Claude-family only).
- normal_implementation (Sol · medium) -> review_codex: Fable 5 · low
  (token-limit fallback Opus 4.8 · low).
- NEW complex_implementation (Sol · high) -> review_codex_complex: Fable 5 ·
  medium (CHANGED from high; token-limit fallback Opus 4.8 · medium).
- fast_iteration (Luna · max, incl. swarm) -> NEW review_codex_fast: Opus 4.8 ·
  medium (token-limit fallback Sonnet 5 · high).
- Forward rule (prose): future max-effort OpenAI impl -> Fable 5 · high review.
- Sol effort-selection addendum (prose): low default / medium mid-slice
  research / high new features / max escalation.

Per the PR #784 doctrine (Fable 5 restored to the plan), the Fable review
primaries are in-plan (subscriptionState included) and auto-selectable — no
approval gate, no Opus substitution condition; the Opus entries on those lanes
exist only as token_limit_fallback (excluded from primary resolution).

Rationale: high Sol-low/medium volume was consuming Fable capacity through
review; Fable is reserved for medium+ pairings. Invariants preserved:
opposite-family review never traded away (all fallbacks Claude-family),
generator != evaluator, no implicit paid escalation.

Adds MODEL_IDS.sonnet (sonnet-5) for the Claude-family token-limit fallback;
drops "review of GPT implementation" from the documentation_review row now that
dedicated review lanes exist. Guard test config/no-hardcoded-volatile passes.


Claude-Session: https://claude.ai/code/session_017LHrkXyMzsQwb9bqr82EFK

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* chore(harness): plan markdown hydration CI fix

* fix(fresh): stabilize hydration builds on clean runners
* chore(harness): plan Aspire generator emission fixes

* fix(cli): emit valid Aspire executable capabilities

* fix(cli): project valid Aspire environment values

* fix(cli): bound Garnet tool restore

* docs(harness): record generator gate evidence
…ed in code (#776)

* feat(agentic): bind opposite-family evaluator routes

* feat(agentic): bind open-model evaluator route

* feat(agentic): promote open evaluator presets
…771)

* chore(tools): add JSR tagline byte-cap gate (docs:tagline:check)

The jsr.io package description is derived from each README's bold tagline by
jsr-set-package-settings.ts and capped at 250 BYTES (Rust String::len) — em-dashes
cost 3 bytes each. Over the cap it is silently truncated at a word boundary, which
is why several published descriptions read as cut mid-sentence today.

This gate extracts the tagline exactly as the release tool does and measures it in
bytes, so an over-cap tagline is caught before publish instead of on jsr.io.

Currently: checked=35 over=16.

Refs #715

* docs(jsr): fit package taglines within byte cap

* ci(docs): gate JSR tagline length

* docs(jsr): align reconciled taglines with formatter
…po-drift CI blocking (#772)

* quality(fresh): type query and middleware boundaries

* quality(sagas-core): type builder and adapter contracts

* quality(streams-core): align environment and schema contracts

* quality(triggers-core): type contract and processor boundaries

* quality(plugin): narrow reserved-name checks without casts

* quality(queue): preserve AMQP channel connection type

* quality(repo): type remaining fixture and plugin boundaries

* ci(quality): block repo-wide TypeScript drift

* fix(quality): reconcile stream and saga contracts after integration

* fix(fresh): resolve Fresh Signals imports on cold builds

* fix(ci): retry transient close-gate API failures

* fix(ci): fall back to public close-gate reads
…nRouter, open models only (#777)

The doctrine already said OpenHands is for cloud-driven runs and that a
local run must use a local adversarial agent for PLAN-EVAL / IMPL-EVAL --
but it named no local transport, so it described a gap and told us to log
it. Name the transport and close the gap.

Local PLAN-EVAL / IMPL-EVAL now runs on Claude Code + OpenRouter
(claude-openrouter profile -> claude-print) with an OPEN model
(minimax/minimax-m3, qwen/qwen3.7-max). An open model is neither
Claude-family nor Codex-family, so it is adversarial to both generators --
the generator-never-evaluates-itself invariant is satisfied more robustly
than by a family swap alone.

OpenHands is unchanged and stays the default automated cloud agent. Its
model rules are inherited verbatim by the local lane: OPEN models only;
closed/paid models (Claude/GPT/Gemini) are PROHIBITED because they burn
paid OpenRouter credit -- cost protection, not a runner implementation
detail. The AGENTS.md CI-gate trigger template is untouched.

Ordinary (non-formal) review remains opposite-family Claude <-> Codex.

Capability is stated PER MODEL, never per transport (drift D-4 amended):
minimax-m3 and qwen3.7-max both return a real reasoning trace and have a
verified agentic turn, so the evaluator can run gates and its effort is
genuine. The zero-reasoning behaviour is specific to GLM 5.2 over
OpenRouter -- a design-lane model -- and must not be restated as a
property of the client, the transport, or the evaluator lane.

The companion routing-policy.ts slice binds the route in data: it adds
qwen to OPENROUTER_MODEL_IDS and makes resolveCanonicalFormalEvaluatorRoute()
throw for anything but Claude + OpenRouter + open_only, so the closed-model
prohibition is enforced in code, not in a comment. The evaluator lane was
the only lane living purely in prose -- which is exactly why an unexamined
assumption could persist in it.

Validation: docs:links 0 broken; agentic:sync-claude:check OK (17 skills);
fmt wrapper zero net-new findings vs baseline (58 vs 61 pre-existing legacy
Markdown drift; net -3, none introduced).


Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…e restoration + OpenCode lane (#799)

chore: reconcile main into beta.10 integration (agentic combo + Fable restoration + OpenCode lane)
…s page, tutorial revamp, stale-reference sweep (#798)

* docs(site): beta.10 refresh — full CLI command reference, agent/skills page, tutorials off manual steps, stale-reference sweep

Derived from the live --help tree of the frozen beta.10 integration state; verified
by links check, internal-wording gate, 12-command accuracy spot check, and
versionless-specifier check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017LHrkXyMzsQwb9bqr82EFK

* docs(site): address beta.10 docs-eval FAIL_FIX for the agentic combo

Merged the beta.10 integration base (agentic combo: @netscript/mcp, the
installed skill bundle, and the `netscript agent` CLI) and reconciled the
docs against the shipped public `netscript` binary.

- ai/mcp.md: drop the stale "NetScript does not host an MCP server" claim;
  distinguish the @netscript/ai/mcp client library from the `netscript agent
  mcp` stdio server and cross-link Agent tooling + the @netscript/mcp
  reference; pin the release specifier.
- reference/cli/commands.md: add the previously-missing `init` and `agent`
  top-level groups (incl. `agent mcp`/`agent init` flags), add the
  `service ref add --project-root` flag, and state the page documents the
  public binary.
- reference/ai/skills.md: document the three installed public skills
  (netscript, netscript-build, netscript-operate) and separate them from the
  loader library surface.
- how-to/author-a-plugin.md: remove the invented `--no-register` flag;
  `--register` defaults on.
- how-to/build-a-durable-chat.md: pin `jsr:@netscript/ai` via releaseSpecifier.
- cli-reference.md: fix the `service add` example (no positional; `--port`).

Verified every command example in the changed files against
packages/cli/bin/netscript.ts; `deno task docs:links` green; Lume build
clean; no internal wording or bare pinnable jsr:@netscript/* on changed lines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(site): precise per-host agent init behavior + complete nested flag coverage

- reference/cli/commands.md: state the per-host `agent init` split precisely
  (Claude host path: .mcp.json + skills under .claude/skills/ + marked
  AGENTS.md section; VS Code host path: .vscode/mcp.json only; --host all
  runs both); add the omitted `service ref remove --project-root` flag.
- Full nested-help sweep of the public `netscript` binary: added the
  remaining omitted flags on documented rows — `generate runtime-schemas`
  (--project-root, --dry-run, --force, --verbose), `generate plugins`
  (--project-root, --dry-run, --verbose), `db list` (--project-root,
  --json), and the `plugin auth` rows (backend set/show --project-root,
  provider set option set, session list --stream-url, session revoke
  --auth-url).
- reference/ai/skills.md: skills and the AGENTS section install only on the
  Claude host path; the VS Code path writes .vscode/mcp.json with no skills.

Verified against packages/cli/bin/netscript.ts recursive --help; docs:links
green; Lume build clean; changed-line greps clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@rickylabs rickylabs added this to the 0.0.1-beta.10 milestone Jul 17, 2026
@rickylabs
rickylabs merged commit 4d438ce into main Jul 17, 2026
27 of 34 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

1 participant