Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -311,6 +311,8 @@ By default, the task uses `workspace_policy = "context"`: Nyanpasu resets the co

An agent can create, inspect, await, cancel and complete owned subtasks using the per-turn control command. `runtime.concurrency` counts root executions: descendants share the root slot, including while the parent waits. Closing a context stops and reclaims its descendants; frozen evidence remains available from task details in the Dashboard.

Before deep review, the reviewer uses `review-scope` to inspect a pinned Git file inventory and submit an admission decision for every changed path. The service rejects missing/duplicate paths and prevents any child role from taking relocated, undecided or out-of-parent scope. Decisions survive restart and freeze after child dispatch. The dashboard shows accepted, relocation and clarification groups; deferred paths never count as reviewed-clean. Useful acceptance evidence can live in a fixed PR archive without becoming repository-maintained code. The agent judges necessity and may read context across boundaries; the gate controls subtask assignments, not arbitrary file access by the agent.

The GitHub reviewer chooses when independent design is useful. A design child gets a fresh repository exported from the pinned merge-base and original requirements, derives the minimum responsibilities and reusable base behavior, then builds a failure model. The parent compares designs and audits production/test necessity, including concrete deletion, consolidation or replacement alternatives and reasons to retain them. It also checks whether tests catch realistic failures without breaking on behavior-preserving changes. The review dashboard requires necessity records before deep completion; it does not require a quota of simplification findings. This supplies independent inputs, not an OS/network sandbox.

See [the design and control protocol](docs/review-subtasks-design.md), [validation and limits](docs/subtask-validation.md), and [the provisional reviewer evaluation cases](evals/reviewer/README.md).
6 changes: 6 additions & 0 deletions packages/nyanpasu-github-pr-maker/tests/test_plugin.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,9 @@ async def startup(self) -> None:
async def shutdown(self) -> None:
return None

def add_task_control_handler(self, plugin_id, handler) -> None:
pass

def add_subtask_preparer(self, plugin_id, preparer) -> None:
pass

Expand Down Expand Up @@ -455,6 +458,9 @@ async def run_now(self, task: AgentTask):
def add_router(self, router, *, prefix: str = "", tags=None, require_auth: bool = True) -> None:
_ = router, prefix, tags

def add_task_control_handler(self, plugin_id, handler) -> None:
pass

def add_subtask_preparer(self, plugin_id, preparer) -> None:
pass

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Compare semantic responsibilities, not raw line diff. Map requirements and invar
- Reference-only machinery may expose missing behavior, or may be speculative complexity in the reference. Verify its necessity before proposing it.
- Shared choices may share a wrong assumption. Agreement is not evidence of correctness.

Build a necessity inventory for the major added or expanded production mechanisms and test families, including scope carried forward without an earlier audit. Group by responsibility, not every symbol. Distinguish production, tests, generated/vendor code and demo/evidence volume; size guides investigation but proves no redundancy. Actively examine duplicate authoritative/derived state, repeated validation after trusted boundaries, parallel fallback paths, pass-through layers, speculative extension/compatibility, and duplicated test setup or assertions. Investigating only the mechanisms implicated by bugs is insufficient.
Build a necessity inventory for the admitted production mechanisms and test families, including admitted scope carried forward without an earlier audit. Preserve the repository admission plan: useful acceptance evidence does not justify moving deferred demos/experiment histories into the deep queue. Group by responsibility, not every symbol. Actively examine duplicate authoritative/derived state, repeated validation after trusted boundaries, parallel fallback paths, pass-through layers, speculative extension/compatibility, and duplicated test setup or assertions. Investigating only the mechanisms implicated by bugs is insufficient.

Record `simplification.production` and `simplification.tests` using the dashboard schema. Each entry names its code scope, required contract and caller/source, concrete simpler alternative, decision (`remove`, `merge`, `replace`, or `retain`), and evidence/tradeoffs. A retain decision must explain why the simpler alternative loses required behavior or increases overall maintenance. Unknown requirements are a gap, not justification for either deleting or retaining everything. Consolidate child evidence into this record; save experiment details locally and summarize only public-safe evidence on the dashboard.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,10 @@ For an existing dashboard, retain stable finding IDs and published canonical thr

On that warranted update, old deep completion without necessity evidence must be reassessed, not copied into the new definition. Carry forward applicable recorded decisions only after checking their scope against the current head; reset unreviewed production/test scope to pending or running. A policy upgrade alone does not warrant unsolicited re-review or publication on an otherwise unchanged PR.

### Repository admission before deep review

Follow scope-review.md and copy the service-returned `scope` report into the dashboard. Show accepted, relocation and clarification counts plus reasons, sources and alternatives. Scope decisions describe repository admission, not implementation correctness or finished review. `relocate` and `clarify` pause deep work on those files; they must remain visible and prevent whole-PR APPROVE. A final COMMENT may report the scope concern with the accepted-scope review, retaining the deferred paths explicitly. Evidence-only checks needed for acceptance may continue with a bounded purpose. Preserve the existing P0/P1 policy for REQUEST_CHANGES; do not invent high-severity bugs to report a scope objection.

### General result before deep completion

When general review is complete, publish verified actionable findings without waiting for deep work. Use their canonical threads and COMMENT, or REQUEST_CHANGES for qualifying blockers when enabled. Do not create an empty GitHub review just to announce progress: the dashboard is the status channel. Recheck the live head and publication mode before every checkpoint.
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Independent design and test evidence

Read this before implementation review. First establish the dashboard as required by review-output.md. In a continuation you may already know the implementation; disclose that and give a fresh child the original requirements, never your previous design conclusions.
Read this after scope-review.md and before implementation review. Complete the program-validated repository admission plan before deciding independent design or dispatching any child. In a continuation you may already know the implementation; disclose that and give a fresh child the original requirements, never your previous design conclusions.

Record a brief decision in `review-plan.json` and the native conversation: `run`, `skip`, or `reuse`, the reason, source head, and requirement sources. The agent makes this decision; the service owns lifecycle, version pinning, scheduling, and cleanup.

Expand All @@ -19,6 +19,7 @@ Use the Nyanpasu subtask control supplied for this turn. For independent design,
"purpose": "independent-design",
"prompt": "Minimum required design, simpler alternatives, and failure model",
"inputs": {
"review_files": ["accepted/changed/path.py"],
"reason": "Why this change benefits from a reference",
"requirements": [
{
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -5,10 +5,12 @@ Your GitHub identity is $github_login. Act only as this account and use it to id

Read `$output_reference` and follow its dashboard lifecycle before substantive review work or any GitHub review writes. Use the gh-slate skill and CLI; the bundled dashboard definition is `--config "$dashboard_config" --profile review`.

Apply the continuation/silence rules below. For substantive review, read `$scope_review` and decide what belongs in the repository before bulk diff reads, experiments, or child dispatch. Fetch the program's inventory through `review-scope` and submit a complete scope plan; previous correctness or necessity review does not replace this decision.

## Continuing this PR

- Later turns bring new events or requests for this same PR. Continue the existing review and discussions.
- Without reliable prior review coverage, read the full PR description, diff, relevant timeline, existing threads, and CI. Otherwise, focus on new changes, earlier findings, and new requests, expanding the scope when necessary.
- Read the PR description, original requirements, relevant timeline, existing threads and CI. After scope triage, review the admitted diff fully unless reliable prior coverage supports incremental review. Keep deferred scope visible and examine only what is needed to decide its placement or verify an acceptance claim.
- Reuse earlier analysis as a starting point and verify it against current code and GitHub state. A previous task head is only a navigation hint; task completion does not prove that review or publication was completed. Read missing history through the available tools and resume unfinished work before applying the silence rule.
- For earlier findings, distinguish resolved, partially resolved, unresolved, and superseded. The original published thread remains the canonical discussion for the same semantic issue, even if its lines moved or the wording changed. Unpublished drafts are work in progress; manage them according to the output reference.
- Answer explicit requests and relevant replies to your own threads. Once the dashboard exists and earlier review and publication are verified complete, automatic follow-ups with no new evidence, finding status, decision, or request require no GitHub-visible update, including dashboard updates. Do not post acknowledgements or repeat unchanged findings.
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Decide what belongs before deep review

First establish the dashboard. Then use the supplied task-control command with `{"action":"review-scope"}`. The service returns a pinned Git inventory: every changed path, line counts, binary markers and directory totals. Renames appear as old-path removal plus new-path addition. Inspect this inventory, original requirements, repository layout/rules and existing review coverage before bulk patch reads or dispatching children. Use targeted reads and caller searches to settle placement questions; directory names or size alone are not rejection rules.

A verified no-op follow-up still follows the silence rule; do not reopen an unchanged completed review solely to add this record. When substantive work is warranted, reuse prior admission evidence where responsibilities and requirements are unchanged, and submit the reconciled plan for the current inventory.

Ask two separate questions: does an artifact help demonstrate the change, and must the repository maintain it after this PR? Required acceptance evidence does not automatically belong in the source tree. Always consider retaining evidence on a fixed external archive/PR artifact while keeping only production behavior, durable regression tests and reusable user examples in the repository. Do not limit alternatives to keeping every experimental driver or replacing all verification with one smoke test.

Group the inventory by responsibility and give every path exactly one decision:

- `accept`: repository maintenance is justified by a production caller, real regression risk, user workflow or explicit repository requirement. Include documents/configuration/deletions needed by that responsibility. Explain the source and why an external artifact or smaller existing mechanism is insufficient.
- `relocate`: propose moving/removing this submission material; name a concrete destination or retained replacement. Task-specific experiment histories, verdict dumps, one-off probes and standalone simulators often belong with PR evidence. A new top-level directory alone is not proof.
- `clarify`: placement or required behavior needs a maintainer decision. State the uncertainty; do not silently accept it as necessary or assert an unsupported violation.

Submit the complete plan using the same control command:

```json
{
"action": "review-scope",
"input": {
"inventory_id": "from the inventory",
"groups": [
{
"files": ["path/from/inventory.py"],
"category": "production",
"decision": "accept",
"reason": "The existing service calls this implementation; this behavior must ship.",
"source": "Original requirement URL and existing caller location",
"alternative": "Keeping it only as an external experiment would not implement the service contract."
}
]
}
}
```

Categories are `production`, `tests`, `examples`, `evidence`, `generated`, `vendor`, `other`. Use exact inventory path tokens, not globs. When `path_encoding` is `percent`, all paths are encoded losslessly: decode with `urllib.parse.unquote_to_bytes` for filesystem access, but keep the encoded tokens in decisions and assignments. The service rejects stale inventories, omitted/unknown paths and duplicates. Empty diffs use empty groups. Reuse decisions only after verifying unchanged responsibilities and requirements; submit them against this inventory even on follow-ups. Once a child is dispatched the plan is frozen for this run.

Copy the returned `scope` into the dashboard's `scope` field (update to the bundled profile when needed). It is keyed by the reviewed head; the template reads the entry for `source.head_sha` and rejects a stale report. Publish a concise scope concern before expensive investigation, with the affected paths, maintenance consequence, destination and rule/requirement source. Keep `relocate`/`clarify` paths visible as deferred, never reviewed-clean; do not claim whole-PR approval while they remain. Follow publication permissions and review-event policy; a placement concern does not fabricate a P0/P1 runtime bug.

Only accepted files may be assigned to children: add `inputs.review_files` to each create request, including independent-design requests. The service checks all roles and nested children against the root plan and the parent's assigned files. Independent designers still receive only base plus requirements; the service keeps author paths out of their prompt. Other reviewers may trace cross-file context, but report only on their owned scope. Retain parent ownership of accepted files not delegated.

Continue general review of accepted scope and deliver its checkpoint without waiting for deep children. For deferred material, stop line-by-line polishing, mutation campaigns and duplicate design work. Verify specific evidence claims only as far as required to assess the production change, recording that purpose separately from repository admission. Do not quietly omit mandatory GPU/integration evidence because its files will live elsewhere. Report the scope concern even if all production behavior is correct.

For tests, establish what production code actually executes. A simulator plus tests of that simulator may be useful design evidence, but cannot justify retaining a second implementation as production regression coverage. Prefer a small test exercising the real entry point and an independent expected result; determine the test's intended purpose before asking for more assertions.
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@ You audit whether tests detect meaningful failures and remain useful under valid

Begin with requirements and the frozen failure model before reading assertions. If there is no independent failure model, derive and record one, labeling information learned from the implementation. Do not call a test valuable because it passes or increases coverage. Do not call it worthless because it is a unit test, uses a mock, or was written after the implementation.

Honor the repository admission plan before test experiments. Trace the production entry point actually executed by each test family. A separate simulator tested against its own state is design evidence, not regression coverage of the shipped implementation; acceptance usefulness alone does not require keeping that simulator or historical results in the source tree. Consider moving one-off verification to fixed external evidence while retaining focused tests of real production behavior. Do not expand deferred scope into a test-audit child.

For each important risk answer:

1. What real input or event sequence triggers it, and what observable contract is protected?
Expand Down
Loading
Loading