Skip to content
Merged
10 changes: 10 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -195,6 +195,8 @@ Dashboard token. Without a webhook secret, that endpoint requires the server tok

Each backend owns its native conversation history. Nyanpasu stores scheduling metadata and session references; the Dashboard reads native messages, reasoning, tool calls, edits, and results without maintaining another conversation database. Session details identify the backend and native session ID. Historical conversations remain readable after changing backends while their runtime and history files remain available.

The session list shows root sessions by default; enable **Show subtask sessions** to include children. Each session has **Conversation**, **Session details**, and **Sub tasks** tabs. Switching tabs preserves reading position, search input, and expanded task groups; the selected tab can be bookmarked. Subtask cards use consistent colors for independent design, module review, and test audit, with text labels and separate status indicators. They include nested children, waiting relationships, result summaries, and downloadable evidence. Open a child conversation and use **Back to parent session** to return to your reading position.

The dashboard frontend is built with Vite+ and managed with pnpm. Use the pnpm
version pinned in `package.json`. During development, use:

Expand Down Expand Up @@ -304,3 +306,11 @@ Plugins that need current external state before execution can register `runtime.
Task merging is opt-in with a `coalesce_key` and requires a registered preparer. Compatible queued tasks with the same key, plugin, and context can merge within `runtime.coalesce_window_seconds`; the preparer owns their domain-specific merge. Ordinary tasks remain separate. Recording and merging happen in one transaction, and a running task cannot receive late merged events. This window does not delay task execution or guarantee that nearby events will share a turn.

By default, the task uses `workspace_policy = "context"`: Nyanpasu resets the context workspace to `workspace.revision` or `workspace.ref`, runs the selected backend there, and keeps that workspace for the next event in the same context. Plugins can opt into `workspace_policy = "event_snapshot"` only when they need a disposable per-event workspace.

### Managed subtasks and independent review

An agent can create, inspect, await, cancel and complete owned subtasks using the per-turn control command. `runtime.concurrency` counts root executions: descendants share the root slot, including while the parent waits. Closing a context stops and reclaims its descendants; frozen evidence remains available from task details in the Dashboard.

The GitHub reviewer chooses when independent design is useful. A design child gets a fresh repository exported from the pinned merge-base, original requirements and a failure-model prompt. The parent compares both designs and audits whether tests catch realistic failures without breaking on behavior-preserving changes. This supplies independent inputs, not an OS/network sandbox.

See [the design and control protocol](docs/review-subtasks-design.md), [validation and limits](docs/subtask-validation.md), and [the provisional reviewer evaluation cases](evals/reviewer/README.md).
695 changes: 695 additions & 0 deletions docs/review-subtasks-design.md

Large diffs are not rendered by default.

41 changes: 41 additions & 0 deletions docs/subtask-validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
# Subtask validation

The tests use real SQLite, local Git repositories, the control CLI in a separate Python process, and real HTTP/browser requests. The model output boundary is controlled with fake execution backends. This validates scheduling, persistence, isolation of supplied inputs and the UI; it does not prove model judgment or sandbox isolation.

Validated behavior includes root capacity with two running children, waiting and continuation, both sides of the child-finish/wait race, restart with preserved workspaces, ownership restrictions, idempotent creation, closing/creation races, stale-generation fencing, cleanup ordering and retry (including creation still running in a Git worker and backend startup failure), expiring controls, frozen evidence and authenticated retrieval after cleanup. Codex cancellation tests check both cancellation during `turn/start` and during execution: an interrupt acknowledgement alone must not permit cleanup.

The reference integration fixture creates divergent head and target branches. It verifies a fresh child at their common base, no PR head object or remote in the exported repository, a usable independent patch baseline unaffected by release export attributes, recorded source identities, immutable target selection under concurrent fetches and changed requirement digests. Template keyword tests are not used as evidence of review quality.

## Directed experiments

Run `uv run python scripts/check-subtask-mutations.py --output /tmp/nyanpasu-subtask-validation` after committing the candidate. The script exports HEAD to a temporary directory, establishes a clean baseline, runs each edit separately and retains full logs. It restores each file and never edits the working repository. The checked substitutions make the experiment reproducible; they are not product assertions or implementation-hash tests.

| Deliberate change | Observed result | Protected behavior |
| ------------------------------------------------------- | ---------------------------------------------------- | ------------------------------------------------------------------------------------- |
| Make a child acquire the root semaphore | Concurrency test fails its bounded child-start check | With one root slot held, children must still start; the unrelated root remains queued |
| Do not consume the persisted wait set | Both early/late child completion cases fail | A completed wait is consumed once and cannot repeatedly resume its parent |
| Start the reference from head instead of merge-base | Reference integration test rejects head-only code | Independent input cannot silently contain the author's implementation |
| Rename the private snapshot-export helper and its calls | Reference and version tests pass | Internal naming can change without changing the observable contract |

The scheduling failure is a deterministic injected deadlock with event-gated tasks, not a timed-out external environment. The head-contamination fixture observes the wrong file before the parent test's bounded wait expires. Do not interpret arbitrary timeouts, collection errors or equivalent mutants as detected defects.

## Recorded local result

On 2026-09-24, Python 3.14.7: **265 pytest tests passed**, Ruff and `ty` passed; frontend checks, type generation and build passed; **6 frontend tests and 15 browser tests passed**. The directed experiments produced the outcomes above. Native model quality was not evaluated.

The 2026-09-25 asynchronous review update reran all **265 pytest tests**, Ruff, `ty`, Markdown formatting, frontend checks and the production build successfully. The 11 offline dashboard cases below also passed. No new model-adherence result is claimed.

The subsequent session-tree update passed **267 pytest tests, 6 frontend tests, and 17 browser tests**, plus Ruff, `ty`, frontend checks and the production build. API cases cover Codex and Claude ownership, nested waits, retained evidence, older task-group pagination, filtering before pagination, and relationships available without reading native history. Browser cases cover subtask navigation, evidence downloads, retained filters and collapsed groups during refresh/failure, and restoration of the exact parent reading position. The navigation case exposed and verified a fix for a lost scroll offset when reopening a session.

The session-tabs update passed **267 pytest tests, 6 frontend tests, and 18 browser tests**. Additional browser coverage checks keyboard tab navigation, bookmarked tabs, empty subtask views, distinct and consistent task-type colors, and preserved conversation DOM, reading position, search input, and collapsed groups across tab switches and background updates. Light, dark, and mobile layouts were inspected separately.

## Commands and limits

For the asynchronous review update, offline `gh-slate render` checks covered 11 dashboard states: initial, preliminary, integrated, deep failure preserving completed general review, skipped deep review, an early blocker, legacy data without stages, and four invalid combinations. The schema rejects preliminary results without completed general review and final approval/comment while deep work remains unfinished. The preview uses the bundled `review.example.json`; no GitHub write is needed. These checks validate the rendered publication contract, not whether a model follows the checkpoint instructions.

- `uv run pytest -q` for backend and plugin behavior.
- `uv run ruff check .`, `uv run ruff format --check .`, and the repository's `ty check --error-on-warning` command.
- `pnpm run types`, `pnpm run check`, `pnpm run test`, `pnpm run build` and `pnpm exec playwright test`.
- Browser fixtures display persisted states without recovering them into actual model runs. This host requires loopback proxy bypass and local Chromium dependencies; those host adaptations are outside the source change.

Not validated by these checks: live Codex/Claude model adherence to the prompts, cross-host execution/failover, model quality/cost on historical PRs, or an OS/network blind sandbox. The capability and base-tree snapshot are not protection from a malicious process with the same OS account. `evals/reviewer/` contains provisional cases and a maintainer grading protocol; no blind evaluation results are claimed. The 2026-09-24 deployment separately verified health, migration, authenticated Dashboard pages and reviewer polling; service readiness is not model-quality evidence.
16 changes: 16 additions & 0 deletions evals/reviewer/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# Reviewer behavior evaluation

The runtime tests validate orchestration. They do not establish that a model identifies elegant designs or valuable tests. These cases are a small, provisional evaluation set; maintainer review and historical PR cases are still required before making quality claims.

Use the same model, effort, environment and original requirements for three conditions: existing review, independent design comparison, and independent failure model plus test audit. Give the independent child only `requirements` and `base`; disclose implementation exposure on continuations. Freeze its artifacts before revealing `head` and `tests`. Keep the grading notes separate from the agent's input. Preserve complete outputs, costs, latency, skipped stages and unavailable checks.

A maintainer judges each finding without seeing which condition produced it: correct reachable defect, useful simplification, acceptable tradeoff, unsupported preference, invalid deletion, or missing evidence. Repeat a subset to assess stability. Do not grade by keyword inclusion, report length, finding count, or the agent's own preference. The four fixtures below are deliberately small; no claim of general effectiveness follows from passing them.

| Case | Files | Question |
| ---------------- | ------------------------------ | ------------------------------------------------------------------ |
| Whitespace | `cases.json`, `whitespace` | Does the test isolate the whitespace-sensitive behavior? |
| Request ordering | `cases.json`, `latest-request` | Does the reviewer construct the actual late-response sequence? |
| Cached size | `cases.json`, `cached-size` | Can duplicate state be removed while preserving observed behavior? |
| Required adapter | `cases.json`, `adapter` | Does the reviewer preserve the real protocol boundary? |

Run code experiments in disposable workspaces. Record a clean baseline, one plausible defect, and a behavior-preserving alternative when relevant. A syntax/import/environment failure is inconclusive. The directed runtime experiments used to validate this implementation are recorded in `docs/subtask-validation.md`; they are separate from model evaluation.
30 changes: 30 additions & 0 deletions evals/reviewer/cases.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
[
{
"id": "whitespace",
"requirements": "Cache keys must distinguish every byte of the user prompt, including leading and trailing spaces.",
"base": "def key(prompt):\n return prompt\n",
"head": "def key(prompt):\n return prompt.strip()\n",
"tests": "def test_inputs_differ():\n assert key('hello ') != key('goodbye')\n"
},
{
"id": "latest-request",
"requirements": "When a user changes the search query, a late response for the previous query must not overwrite the latest results.",
"base": "let rows = [];\nasync function search(query) { rows = await fetchRows(query); }\n",
"head": "let rows = []; let current = '';\nasync function search(query) { current = query; const result = await fetchRows(query); rows = result; }\n",
"tests": "// fetchRows always resolves immediately.\nawait search('old'); await search('new'); expect(rows).toEqual(['new']);\n"
},
{
"id": "cached-size",
"requirements": "Queue supports push, pop and size. There is no cached-size persistence or API contract. Python len(list) has constant time for this workload.",
"base": "class Queue:\n def __init__(self): self.items = []\n def push(self, item): self.items.append(item)\n def pop(self): return self.items.pop(0)\n",
"head": "class Queue:\n def __init__(self): self.items = []; self.count = 0\n def push(self, item): self.items.append(item); self.count += 1\n def pop(self):\n value = self.items.pop(0)\n self.count -= 1\n return value\n def size(self): return self.count\n",
"tests": "def test_size():\n q = Queue(); q.push('a'); q.push('b'); q.pop(); assert q.size() == 1\n"
},
{
"id": "adapter",
"requirements": "All callers use seconds. The external transport API uses milliseconds. The adapter is the only code allowed to depend on that wire convention.",
"base": "# A new timeout option is needed on the existing transport boundary.\n",
"head": "class Transport:\n def __init__(self, wire): self.wire = wire\n def send(self, timeout_seconds): return self.wire.send(timeout_ms=timeout_seconds * 1000)\n",
"tests": "def test_wire_units():\n wire = Mock(); Transport(wire).send(0.25); wire.send.assert_called_once_with(timeout_ms=250)\n"
}
]
8 changes: 8 additions & 0 deletions evals/reviewer/grading-notes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Provisional grading notes — withhold from the reviewer

- **Whitespace:** the test changes both the text and its trailing whitespace, so it misses `strip()`. Compare `'hello'` with `'hello '` only. A direct string expectation is an independent oracle here because the requirement explicitly preserves bytes. Do not replace the check with a hash of the implementation.
- **Latest request:** sequential immediate responses never create the reported race. Start old and new requests, resolve new first, then old; assert that new results remain. A mock transport is useful if real request coordination executes. Merely asserting that a generation counter increments would not prove the behavior.
- **Cached size:** a derived `len(items)` can remove the maintained count and synchronization points. Preserve the public `size()` contract and exercise mixed pushes and pops. The existing size test is useful under that refactor; do not delete it simply because a simpler implementation exists.
- **Adapter:** the wrapper has a required protocol responsibility. The mock assertion targets the actual external call contract and its units. Removing the adapter or rejecting the test just because it asserts a call would be a false positive. A separate real transport check may still be needed for transport compatibility.

These expectations are author-proposed, not maintainer-confirmed results. An alternative finding requires a reachable trigger and independent evidence. Mark an unavailable execution boundary honestly; do not infer a pass from plausible code.
Loading
Loading