Skip to content

factory: state-machine workflow executor (factory.run/status/stop/resume, guarded transitions, re-entry) - #2401

Open
sethkarten wants to merge 13 commits into
swarm/swarm-spec-harness-entryfrom
swarm/swarm-dag-executor
Open

sethkarten wants to merge 13 commits into
swarm/swarm-spec-harness-entryfrom
swarm/swarm-dag-executor

Conversation

@sethkarten

@sethkarten sethkarten commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

PR G of the factory DAG feature: the executor. Builds on PR #2397 (factory entry kind, validate_factory_spec / canonicalize_factory_spec / topological_order).

  • Python kernel side (prime-agent-runtime/src/rlm/factory.py + rlm/__init__.py): a FactoryExecutor exposed as the rlm.factory namespace — await rlm.factory.run(spec_id, name=?), rlm.factory.status(run_id), rlm.factory.stop(run_id), rlm.factory.resume(run_id).
  • TS host side: a "factory.progress" host handler in agent-session.ts that injects one quiet steer-lane notice per run milestone (mirroring the bash.completed path), plus a hint sentence in the model prompt.

State machine model

The factory specification is now a state machine stored in arguments["machine"]. The original DAG form (arguments["dag"]) stays as sugar and compiles to machine form at canonicalization: a compiled dag is a restricted machine — each state entered at most once, forward-only, with guard-less transitions (one per effective dependency edge) and max_entries: 1.

{
  "run": {
    "budget_ms": 600000, "failure_policy": "escalate", // fail_fast | continue | escalate (default)
    "max_parallel": 8,                                   // 1..64 (default 8)
    "max_transitions": 100                               // default 10 * states, hard cap 10000
  },
  "states": [                       // 1..1024, unique slug ids, >= 1 entry state
    {
      "id": "reviewing",
      "entry": true,                 // admission enters every entry state
      "max_entries": 4,              // bounded re-entry; further transitions are blocked
      "subagent": "reviewer",        // required unless the state is a wait state (same forms as V1)
      "lifecycle": "task",           // task | resident
      "inputs":  [{ "name": "pr_url", "type": "text", "from": "entry.pr_url" }], // binds the LATEST settle of the source
      "outputs": [{ "name": "verdict", "type": "json" }],
      "budget_ms": 300000, "retries": 2, "failure_policy": "escalate",
      "foreach": { "over": "items", "max": 64 },   // expands per entry
      "wait": { "kind": "path", "target": "/tmp/x", "timeout_ms": 5000 }  // wait states only
    }
  ],
  "transitions": [
    { "from": "reviewing", "to": "fixing", "on": "settled",
      "when": { "output": "verdict", "path": "approved", "op": "eq", "value": false } }
  ]
}
  • Executor semantics: ALL transitions whose guards pass fire from one settle (fan-out is legal; a guard-less transition always passes; switch patterns use mutually exclusive guards). A transition fires only if the target has entries left under max_entries; blocked transitions are recorded as transition_blocked ledger events and the machine continues. Re-entry re-binds inputs from the current latest settles of the input sources. Self-loops and back-edge cycles are legal — the V1 acyclicity check is removed entirely. Completion is quiescence: no state in flight (running/pending/waiting) and no firable transition from any unconsumed settle. max_transitions exceeded pauses the run once with a budget_exceeded milestone (resume-able).
  • Guards (when, over the from state's latest settle output): ops eq ne gt gte lt lte exists contains; path resolves dotted paths inside json ports; numeric ops require numeric values, contains requires a list, eq/ne scalars.
  • Wait states (wait block): no subagent, declared outputs, retries, or foreach, and never resident; entering registers an rlm.watch.path / rlm.watch.agent host watch and settles the implicit json port event: {"detail", "timed_out"} — from the watch outcome or its timeout through the injectable clock.
  • Unchanged from V1: budgets (per entry: admission → settlement), retries (per entry attempt), foreach (per entry), state failure policies (fail_fast cancels in-flight, escalate pauses, continue finishes), resident teardown by stop(), notices, and the ledger's arrived/shown/read stages.

This PR reworks the executor onto the machine form: run() canonicalizes the stored spec (dag sugar compiles), admission enters the entry states, and the control loop evaluates transitions on every settle. All 32 V1 dag executor tests run unchanged through the compiler path (43 tests total with the machine suites). New ledger event kinds: state_entry (with entry index), transition_fired, transition_blocked (plus wait_settled accepted by the replay checker for legacy ledgers); status() adds per-state entries_used / max_entries, per-entry breakdowns, and a transitions_fired usage count over a 200-event window.

Review-fix round: wait states are gated at the validator (the rlm.watch.* host handlers arrive with the communication series #2351/#2356), so this PR removes the executor's wait registration/timeout/settle/cancel paths entirely — quiescence, re-entry, and fan-out are unchanged. A max_transitions pause mid-settle now records the transition index where it landed and a resume continues after it (previously the settle re-fired already-fired transitions, duplicating spawned work); eq/ne guards compare JSON-strictly (a bool never equals a number, numbers compare numerically); contains requires a non-empty list value at validation and is defensively false at runtime; the max_transitions pause fires its own max_transitions_exceeded milestone kind (no collision with the run-budget budget_exceeded — both notice-able); the control-loop stall names the pending entry whose input source never settled (with a test); EVENT_WINDOW grows 50 → 200 (the pr-manager happy path is ~43 events before any retry); refinement validateEdit accepts machine-form factory edits (either arguments.dag or arguments.machine object, never both). Machine inputs may declare "optional": true to bind a null sentinel instead of waiting when their source state never settled — the enabler for the closed pr-manager review/fix loop in the eval PR (compiled dags never emit it, so dag semantics stay V1-exact).

Spec: Swarm DAGs: declarative orchestration in Continual Harness (sections "Executor" and "Runtime contract").

Executor semantics

Ownership split. The RLM supervisor owns the children (admission via rlm.spawn, settlement via rlm.collect, cancellation via rlm.delete_subagent); the executor owns run state in kernel memory. Every host call resolves through the module-level rlm functions / host_request at call time, so tests patch rlm.host_request directly.

Dry run = write time + run time. create_factory validates the graph at write time. run() re-validates and canonicalizes the stored dag, then resolves every node's subagent reference — a string is a harness subagent entry id or title (entry content becomes the prompt template; optional metadata.model / metadata.thinking carry spawn settings), an inline object uses its own fields. All reference failures are reported together in one ValueError and nothing starts. The resolved node count and max_parallel are reported in the result. Actual admission limits (concurrency, recursion depth, provider rate limits) are enforced at spawn time through the backoff path; that split is the V1 "runtime tree limits" dry run.

Nonblocking admission. run() initializes per-state state, creates an entry for every entry state (a state_entry ledger event each), then prepares and admits instances up to max_parallel (respecting foreach expansion), records handles, and returns {"run_id", "spec_id", "name", "nodes", "max_parallel", "started", "pending"}. A kernel asyncio task (the same loop.create_task pattern as the job-watch poller, guarded so a dead bridge fails the run gracefully instead of wedging) continues the work; the calling model turn ends immediately.

Control loop (iterative, no recursion). Polls in-flight instances with await rlm.collect(ids, timeout_ms=2000). For each settled instance: capture the answer, duration_ms, mark done, append a ledger event; when an entry's last instance settles, the entry settles and captures the state's declared output ports (json ports parsed as in binding, text ports the captured string). Every settle is then evaluated once: ALL guard-passing outgoing transitions fire (fan-out is legal), each fire enters its target unless it is out of max_entries (recorded as transition_blocked and the machine continues), and re-entry re-binds inputs from the current latest settles. Completion is quiescence: no entry in flight (pending/running) and no unevaluated settle. (Wait states are specified but gated: the rlm.watch.* host handlers arrive with the communication series #2351/#2356, so the validator rejects wait blocks and this PR ships no wait execution paths.) Yields once per iteration so an instantly-settling host cannot hot-spin. Tested at 100-node chain and 1000-node wide-fan scale against a fake host.

Input binding. For each input {name, type, from: "<node>.<output>"}: text ports take the upstream captured answer; json ports parse the upstream text — preferring the trailing fenced ```json block whose object contains the output name, else the whole text, else the node errors. Rendering replaces {input_name} placeholders in a single pass; inputs without a placeholder are appended in a trailing ## Inputs section, so no bound value is ever dropped. Binding failures never retry (a deterministic binding error would recur on every re-render); the node fails and its failure_policy applies.

foreach. When the over json input resolves to a list of K items, the node spawns K instances (each item bound as that input's value), K clamped to foreach.max. A foreach node is done when all instances settle; its captured answer is its instances' answers joined with blank lines. A node fails at its FIRST permanently failed instance — it does not wait for its remaining instances — so fail_fast cancels in-flight siblings while they are still running, and a failure that settles before its successes can never strand the node in a mixed terminal state.

Retries and failure policies. A child failure (collect error status/field) re-spawns the same rendered prompt while attempts <= retries; then the node failure_policy applies: fail_fast → cancel every running child of the run via rlm.delete_subagent and mark the run failed; continue → mark the node error and keep going (dependents without a data edge still run; data dependents fail at binding); escalate (default) → mark the node error and pause the run, spawn nothing further until resume().

Budgets (injectable clock). Wall-clock budgets measure admission to settlement. Per node: if the spawn-to-settle span exceeds budget_ms, the attempt is treated as failed — the budget is spent, so no retry; the failure_policy applies. Run-level: elapsed time past run.budget_ms pauses new spawns (children already in flight keep running) with a budget_exceeded milestone; the budget milestone reports once per run, and resuming after it is an explicit operator decision.

Rate limits. If spawn admission fails with a rate-limit/429-style error, the executor retries with exponential backoff — doubling delays capped at 60s, at most 5 admissions per attempt — then fails the node through its failure_policy; the loop never crashes. In the run() admission phase (and in resume(), exactly the same way) a rate limit does not sleep inside the calling turn: the node stays pending and the control loop retries it with backoff (keeps both calls nonblocking).

stop() and races. stop() sets a transitional stopping state before its first await, so the control loop can neither admit new children nor finalize the run while the cancellations are in flight; the loop re-checks state before completion, _finalize only runs while the run is running, and the executor-error path never overwrites a concurrent stop. A stopped run therefore stays stopped (never done + a spurious finished notice). Repeated stop() is idempotent and appends only one run_stopped ledger event.

Event ledger. Kernel memory, per run: spawned/attempts, settled (with duration_ms), answer captured, retries, errors, and run milestones. Stages follow the spec: arrived (child answer settled and captured), shown (milestone notice injected), delivered (parent read via status() — marked on each call). status() returns state states (with per-state entries_used/max_entries, per-entry breakdowns, and per-instance details), the trailing 50 events, elapsed_ms, and usage (spawns/settled/tool uses/running/transitions_fired). status/stop/resume raise ValueError on an unknown run id.

Milestone notices. One quiet runtime notice per milestone kind per run (finished / failed / paused / budget_exceeded), never per node, sent via host_request("factory.progress", {run_id, kind, node?, detail}). The TS handler (createFactoryProgressHostHandler + createFactoryProgressMessage) validates the payload and injects a quiet custom message (customType factory_progress_notice, steer lane, queueIfBusy/resumeIfIdle, suppressAutonomousContinuation, mirroring bash.completed) that wakes the parent with [factory-progress run:<id>] finished|failed|paused|budget-exceeded: <detail>. A dead bridge cannot be told; the ledger keeps the milestone and status() still surfaces it.

Resident nodes (V1 wake sources). A resident node is admitted like any other node and then stays alive under the parent session; its wake source is its own subagent prompt/tooling (watches, heartbeats, or continuous work inside the child) — no structural dry-run rule in V1. A run's declarative work completes when all non-resident nodes settle; resident children keep running (they cannot outlive the parent session) until rlm.factory.stop(run_id) tears them down.

Known limitations (documented deliberately)

  • Runs do not survive a kernel restart. Run state lives in the executor's kernel-memory registry; after a restart the registry is gone. Children are supervisor-owned and keep running — rlm.list_subagents() can still see and stop them.
  • Answer-preview cap. Answers are captured from rlm.collect previews, which the host caps at 160 characters (compactRlmText); the executor additionally caps at 200 defensively. Input binding and downstream prompts work on these previews; full child outputs stay in the child's own session.
  • A paused run does not observe child settlements until resume() (completed children keep their results for a later collect).
  • Run completion with any node error reports state failed (continue policy finishes the graph, but the run is marked failed so the parent sees the error); stop() reports state: "stopped" and lists every non-terminal node as cancelled.
  • Node budgets are settlement-only. A child that never settles is never budget-failed, so the loop keeps polling it indefinitely (V1 limitation); the natural follow-up is a per-instance watchdog. Run budgets share this: they pause new spawns, never kill in-flight children.
  • Cancellation shortcut (V1). If delete_subagent fails, the instance is still reported cancelled in the ledger and a cancelled event is recorded next to cancel_failed (the child keeps running supervisor-side; the executor treats its slot as released). Instances of cancelled entries are marked cancelled with their own ledger event — including never-admitted (rate-limit-deferred) ones and admissions that land after stop() (the returned child is deleted, not registered running).

Tests

  • prime-agent-runtime/test/test_factory_executor.py (49 tests, deterministic, no live model; patched rlm.host_request async fake routing by request type with per-child-id scripted outcomes, collect/delete suspension gates, injected fake clock and sleeps): dry-run rejection (invalid dag via a bypassed store entry, missing references listing all failures, unknown spec), admission starting exactly the in-degree-0 nodes with max_parallel respected, propagation with exact rendered prompts (text, fenced-json, whole-text json, bad json → node error without spawning, unplaced inputs appended, two-parent fan-in in one collect batch, answer-capture cap slice), foreach expansion + clamp + zero items, all three failure policies (fail_fast cascade cancels running children; continue finishes remaining nodes; escalate pauses + notices + resume() continues) plus mixed-instance foreach per policy (fail_fast cancels in-flight siblings; escalate pauses; continue finishes with node error), retries until attempts exhausted, node and run budgets via the fake clock, stop() cascade + states + idempotency, a gated stop-race test proving a stop mid-collect keeps stopped, finalizes nothing, and admits no new children, 429 admission deferral → backoff → success and backoff exhaustion → node error (including the resume() deferral path), 100-node chain and 1000-node fan scale under a generous wall bound, status delivered-marking + unknown-run ValueError, resident node spawn/stay/stop, plus stop-window admission races (a spawn landing after stop() is deleted and never registered; a backoff waking in a stopped run never retries), repeated-pause milestones in the ledger, multi-fence json binding, shown stages surviving a status() read, re-entered-state stop cancelling an in-flight earlier entry, and the cancel_failed/cancelled pairing on failed deletes.
  • packages/coding-agent/test/factory-executor.test.ts (13 tests, mirroring the async-bash completion test): bracket-grammar notice content, convertToLlm pass-through, startsAgentRun wake, payload validation and forwarding, invalid payload rejection.
  • prime-agent-runtime: uv run python -m unittest discover -s test — 483 tests, only the pre-existing test_bash failure(s) reproduced on the clean base branch (test_windows_without_bash_raises_teaching_error; test_unconfigured_journal_stays_permissive also fails on some machines).
  • packages/coding-agent: npx vitest --run test/factory-executor.test.ts test/refinement.test.ts — 105 passed. Repo root: npx tsgo --noEmit and npx biome check on changed files — clean.

Stacked: base branch is swarm/swarm-spec-harness-entry (#2397); merge that first.

Draft only - review withheld per workflow.


Note

High Risk
Introduces async orchestration over subagent spawn/collect/delete with in-memory run state, race-prone stop/resume paths, and parent-session message injection—core delegation and session continuity are affected.

Overview
Adds the state-machine factory executor behind await rlm.factory.run/status/stop/resume: stored dag or machine specs are canonicalized, entry states are admitted via rlm.spawn, and a background control loop settles children, evaluates guarded transitions (fan-out, re-entry, max_entries blocks), and applies retries, budgets, and fail_fast/continue/escalate policies until quiescence or pause/stop.

Run milestones reach the parent through a new factory.progress host path that injects quiet factory_progress_notice custom messages (treated as agent-run boundaries). Harness/refinement docs and validation now accept either arguments.dag or arguments.machine (not both); machine validation gains optional inputs and stricter contains guards. Wait-state execution remains rejected until rlm.watch.* lands.

Large test suites cover executor semantics (Python) and progress messaging (TypeScript).

Reviewed by Cursor Bugbot for commit 6dfd619. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add FactoryExecutor state-machine workflow runner with rlm.factory run/status/stop/resume API

  • Adds FactoryExecutor in factory.py. It runs canonical factory machines through RLM supervisor children, tracks runs in an in-memory registry, and evaluates guarded transitions, retries, failure policies, budgets, re-entry, and cancellation in an async control loop
  • Exposes rlm.factory.run, status, stop, and resume through the new _RLMFactoryNamespace in init.py, backed by module-level wrappers on a lazy default executor
  • Adds factory.progress host handling in messages.ts, agent-session.ts, and rlm-runtime.ts so run milestones appear as agent messages that start a new agent run
  • Accepts either machine or dag factory arguments in refinement edits (refinement.ts) and rejects both, neither, or non-object forms; tightens validation of boolean input optional flags and non-empty contains-guard values
  • Updates harness prompts to advertise the new run/status/stop/resume controls
  • Behavioral Change: factory runs live in-process in the default executor registry; stopping a run deletes supervisor children and failed deletions leave children supervisor-owned; runs with a wait state (deferred inputs) are not yet supported and fail validation
📊 Macroscope summarized 6dfd619. 10 files reviewed, 8 issues evaluated, 2 issues filtered, 6 comments posted

🗂️ Filtered Issues

packages/coding-agent/src/core/prompts/rlm.ts — 0 comments posted, 1 evaluated, 1 filtered
  • line 49: The new prompt omits await for rlm.factory.status(run_id), rlm.factory.stop(run_id), and rlm.factory.resume(run_id), although all three namespace methods are async. Following this instruction in the Python REPL merely creates unexecuted coroutine objects, so the model cannot actually inspect, stop, or resume a factory run. [ Cross-file consolidated ]
packages/coding-agent/src/core/refinement/refinement.ts — 0 comments posted, 1 evaluated, 1 filtered
  • line 721: The newly added factory usage hint omits await for both rlm.factory.status(run_id) and rlm.factory.stop(run_id), although those methods are declared async in the runtime. A model following this prompt merely creates unexecuted coroutine objects, so it neither observes a run nor stops it; in particular, a resident or stuck factory run continues consuming children. [ Cross-file consolidated ]

@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Prime Agent performance — completed

PR 6dfd619c compared with main 08ff1b2e.

Overall: 0 regressed · 0 improved · 42 no clear change.

Metric Main This PR Change
Cold startup 788.9 ms 778.6 ms ≈ -10.3 ms (-1.31%)
Warm startup 589.3 ms 593.2 ms ≈ +3.9 ms (+0.66%)
Installation 6.42 s 6.10 s ≈ -0.32 s (-5.03%)
Compressed release artifacts 73.14 MB 73.23 MB ≈ +0.09 MB (+0.13%)
Installed footprint 595.63 MB 596.13 MB ≈ +0.50 MB (+0.08%)
Idle memory, summed RSS 652.20 MB 652.27 MB ≈ +0.06 MB (+0.01%)

Python runtime

Metric Main This PR Change
Python kernel startup 35.7 ms 34.8 ms ≈ -0.9 ms (-2.50%)
Python cell round trip 0.095 ms 0.074 ms ≈ -0.022 ms (-22.69%)
Empty bash command 2.2 ms 1.9 ms ≈ -0.3 ms (-13.54%)
Bash git status 2.8 ms 2.6 ms ≈ -0.3 ms (-9.42%)
Bash 32 KiB output 2.2 ms 2.1 ms ≈ -0.2 ms (-6.85%)
35 cells / 9 shell calls 29.1 ms 24.7 ms ≈ -4.4 ms (-15.08%)
Python interrupt to done 0.588 ms 0.544 ms ≈ -0.044 ms (-7.48%)
Python state snapshot 10.0 ms 10.0 ms ≈ -0.0085 ms (-0.09%)
Python state restore 138.1 ms 132.4 ms ≈ -5.7 ms (-4.16%)
Python idle RSS 21.29 MB 21.57 MB ≈ +0.28 MB (+1.32%)
Python RSS after pandas workload 75.59 MB 75.88 MB ≈ +0.29 MB (+0.39%)

Session transport

Metric Main This PR Change
Full-history transfers per warm session switch 1.00 transfers 1.00 transfers ≈ +0.00 transfers (+0.00%)
Private frame decode, 32 MiB in 8 KiB chunks 15.9 ms 16.8 ms ≈ +0.8 ms (+5.20%)

UI interactions

Metric Main This PR Change
Resume large session (cold) 1,955.8 ms 1,621.2 ms ≈ -334.7 ms (-17.11%)
CPU, resume large session 2,300.0 ms 2,050.0 ms ≈ -250.0 ms (-10.87%)
Switch into large session 1,691.2 ms 1,466.9 ms ≈ -224.3 ms (-13.26%)
CPU, switch into large session 1,710.0 ms 1,640.0 ms ≈ -70.0 ms (-4.09%)
Open agents view from a session 140.1 ms 140.8 ms ≈ +0.7 ms (+0.49%)
CPU, open agents view 100.0 ms 80.0 ms ≈ -20.0 ms (-20.00%)
Full agents roster, many sessions 4.03 s 4.03 s ≈ -0.0022 s (-0.05%)
CPU, full agents roster 0.92 s 0.95 s ≈ +0.03 s (+3.26%)
Open another session from agents view 1,976.5 ms 1,924.0 ms ≈ -52.5 ms (-2.66%)
CPU, open from agents view 1,140.0 ms 1,110.0 ms ≈ -30.0 ms (-2.63%)
Reopen resident large session 211.5 ms 223.7 ms ≈ +12.2 ms (+5.76%)
CPU, reopen resident session 230.0 ms 220.0 ms ≈ -10.0 ms (-4.35%)
Open subagent session at depth 6 17,552.2 ms 17,690.3 ms ≈ +138.1 ms (+0.79%)
CPU, open subagent at depth 6 4,550.0 ms 4,610.0 ms ≈ +60.0 ms (+1.32%)
Open chain parent from agents view 3,020.3 ms 3,120.1 ms ≈ +99.8 ms (+3.30%)
CPU, open chain parent 1,470.0 ms 1,430.0 ms ≈ -40.0 ms (-2.72%)
Scheduled catalog, first request 422.6 ms 447.3 ms ≈ +24.7 ms (+5.85%)
CPU, scheduled catalog 810.0 ms 880.0 ms ≈ +70.0 ms (+8.64%)
Scheduled catalog, repeated request 0.5 ms 0.7 ms ≈ +0.2 ms (+42.05%)
CPU, repeated catalog 0.0 ms 0.0 ms ≈ +0.0 ms (N/A)
Cold worker with three catalog scans 458.0 ms 433.3 ms ≈ -24.7 ms (-5.40%)
CPU, cold worker and scans 480.0 ms 510.0 ms ≈ +30.0 ms (+6.25%)
UI memory after interactions 1,728.38 MB 1,740.07 MB ≈ +11.69 MB (+0.68%)

Sandbox cost: ~$0.1107 — no inference calls.
Run, logs, and downloadable raw results

Methodology and samples

Main resolved at 2026-09-28T23:14:46.641848+00:00. Harness 08ff1b2e.
Linux x64, 4 vCPU, 8 GB RAM, 20 GB disk; region us.
Image: node:24-bookworm@sha256:be23f54a88d34e8824c741b19b91064094f92c1c97b194144bfc8b50d67258e2.
Stock tools, skills, daemon, and Python bootstrap enabled; fresh homes and a fixed Git fixture.
Onboarding is dismissed; the editor starts without a selected model or submitted prompt.
Medians shown. Arrows require a 20% timing/memory change plus absolute floors and IQR.
These practical noise floors are not a statistical significance test.
Cold means stopped Prime processes; OS filesystem caches are not flushed.
No model requests or credentials. Installation excludes build/setup time.
Installer tarballs use loopback; npm/Python downloads use the network with fresh caches.
Artifact size counts release tarballs; footprint after first use includes registry packages.
MB is decimal. Summed RSS can double-count shared pages; PSS is recorded when available.
Provisioning, setup, and build durations are recorded separately in the raw results.
Kernel probes use the installed JSONL runtime, outside the TUI/TypeScript host.
Per trial: 50 Python cells, 5 calls per shell case, and one 35-cell mix (9 git status calls).
Cell/shell values are batch means; other runtime timings are single operations.
State fixture: a 10,000-row × 8-column integer DataFrame and a 10,000-integer list.
Restore runs in a fresh kernel, including pandas imports; kernel startup is excluded.
Kernel RSS covers the isolated Python process; loaded RSS follows the pandas workload.
Transport benches run node against the prepared source build, outside the installed home.
The switch benchmark drives one warm switch into a 48k-entry session through a real
daemon and counts full-history crossings: streamed replacement snapshots, inline
replacements, and full-history refetch responses.
Frame decode times one 32 MiB private frame, snapshot-chunk header, pushed in
8 KiB chunks; the wire shape of multi-MB frames on the daemon-worker channels.
UI trials use a fresh fixture set: 194 top-level sessions including one ~40 MB transcript,
40 ledger fan-out children, and a 6-deep subagent chain (~46 spawn edges).
Large fixtures hold 1,999 complete triples (~5 MB JSONL); medium 119; subagents 399 each.
Interactions: cold --resume of a large session, warm /resume switch, left-arrow to agents view,
roster settle with many saved sessions, search-and-open of another large session,
reattaching to that resident session, opening the chain parent, and drilling to depth 6.
Readiness is the rendered transcript tail plus a confirmed editor echo.
CPU metrics sum utime+stime across the whole benchmark-user process tree per interaction.
UI memory sums RSS after the interactions; PTY byte counts are in the raw results.
A separate catalog fixture has 2,300 sessions, 2,298 edges, and 13 paused scheduled-job owners.
Catalog timings cover first/repeated reads and cold worker creation under three pending scans.
All expected jobs and owner metadata are checked; worker readiness excludes TUI rendering.
Costs estimate full sandbox lifetimes at configured rates, including setup and build.
Budget target: $1; not a billing cap. Performance changes are informational.
Failed or incomplete execution fails the workflow; saved artifacts remain available.
Each side stops a phase after 2 identical consecutive failures.
Skipped trials are not attempted samples. Warm startup requires a successful cold launch.

Metric Main successful/attempted PR successful/attempted Main spread PR spread
Cold startup 10/10 10/10 IQR 197.4 ms IQR 171.8 ms
Warm startup 10/10 10/10 IQR 45.1 ms IQR 68.2 ms
Installation 3/3 3/3 range 1.36 s range 0.92 s
Compressed release artifacts 1/1 1/1 — —
Installed footprint 1/1 1/1 — —
Idle memory, summed RSS 10/10 10/10 IQR 6.29 MB IQR 4.89 MB
Python kernel startup 10/10 10/10 IQR 6.7 ms IQR 2.7 ms
Python cell round trip 10/10 10/10 IQR 0.042 ms IQR 0.015 ms
Empty bash command 10/10 10/10 IQR 0.4 ms IQR 0.1 ms
Bash git status 10/10 10/10 IQR 0.4 ms IQR 0.4 ms
Bash 32 KiB output 10/10 10/10 IQR 0.6 ms IQR 0.3 ms
35 cells / 9 shell calls 10/10 10/10 IQR 5.6 ms IQR 0.8 ms
Python interrupt to done 10/10 10/10 IQR 0.102 ms IQR 0.068 ms
Python state snapshot 10/10 10/10 IQR 0.3 ms IQR 0.9 ms
Python state restore 10/10 10/10 IQR 9.9 ms IQR 3.3 ms
Python idle RSS 10/10 10/10 IQR 0.09 MB IQR 0.13 MB
Python RSS after pandas workload 10/10 10/10 IQR 0.18 MB IQR 0.18 MB
Full-history transfers per warm session switch 10/10 10/10 IQR 0.00 transfers IQR 0.00 transfers
Private frame decode, 32 MiB in 8 KiB chunks 10/10 10/10 IQR 2.8 ms IQR 3.5 ms
Resume large session (cold) 3/3 3/3 range 338.7 ms range 301.7 ms
CPU, resume large session 3/3 3/3 range 110.0 ms range 50.0 ms
Switch into large session 3/3 3/3 range 388.3 ms range 71.6 ms
CPU, switch into large session 3/3 3/3 range 330.0 ms range 170.0 ms
Open agents view from a session 3/3 3/3 range 7.8 ms range 1.1 ms
CPU, open agents view 3/3 3/3 range 30.0 ms range 60.0 ms
Full agents roster, many sessions 3/3 3/3 range 0.0025 s range 0.41 s
CPU, full agents roster 3/3 3/3 range 0.03 s range 0.21 s
Open another session from agents view 3/3 3/3 range 263.8 ms range 339.4 ms
CPU, open from agents view 3/3 3/3 range 170.0 ms range 150.0 ms
Reopen resident large session 3/3 3/3 range 37.0 ms range 19.6 ms
CPU, reopen resident session 3/3 3/3 range 40.0 ms range 40.0 ms
Open subagent session at depth 6 3/3 3/3 range 86.3 ms range 201.7 ms
CPU, open subagent at depth 6 3/3 3/3 range 90.0 ms range 150.0 ms
Open chain parent from agents view 3/3 3/3 range 179.7 ms range 218.8 ms
CPU, open chain parent 3/3 3/3 range 80.0 ms range 110.0 ms
Scheduled catalog, first request 3/3 3/3 range 37.2 ms range 86.0 ms
CPU, scheduled catalog 3/3 3/3 range 100.0 ms range 120.0 ms
Scheduled catalog, repeated request 3/3 3/3 range 0.1 ms range 0.2 ms
CPU, repeated catalog 3/3 3/3 range 0.0 ms range 0.0 ms
Cold worker with three catalog scans 3/3 3/3 range 98.0 ms range 117.6 ms
CPU, cold worker and scans 3/3 3/3 range 100.0 ms range 50.0 ms
UI memory after interactions 3/3 3/3 range 98.09 MB range 11.54 MB

sethkarten added a commit that referenced this pull request Sep 16, 2026
…p race)

Address the PR #2401 review findings:

- foreach mixed-instance outcomes (finding 1): a node now fails at its
  FIRST permanently failed instance instead of waiting for all instances
  to settle. Previously a failure that settled before its siblings left
  the node stuck in running with every instance terminal and the failure
  policy never applied (the run ended via the stall guard); fail_fast
  could also never cancel in-flight siblings. The policy guard keeps the
  second and later permanent failures no-ops.
- stop() race (finding 2): stop() sets a transitional "stopping" state
  before its first await; the control loop re-checks state before the
  completion check, _finalize only runs while the run is running, the
  node-failure policy records node errors but never overwrites a state
  another transition owns, and the executor-error path no longer
  overwrites a concurrent stop with failed. A stopped run now stays
  stopped: no new children admitted during the stop window, no spurious
  done + finished notice. Repeated stop() is idempotent (finding 7) and
  appends only one run_stopped ledger event.
- resume() defers rate-limited admissions to the control loop with
  backoff exactly like run() (finding 3), so it never sleeps in the
  calling model turn.
- fail_fast milestone detail now counts cancelled in-flight children.
- tests: the host fake is now an async callable patched in directly
  (AsyncMock does not await async side effects) with per-child-id
  scripted outcomes and collect/delete suspension gates (finding 6) so
  races and mixed instances reach real paths; new tests cover mixed-
  instance foreach for every policy, the gated stop race, repeated stop,
  two-parent fan-in in one collect batch (finding 8), the
  ANSWER_CAPTURE_CAP slice (finding 9), and the resume deferral path.
@sethkarten

Copy link
Copy Markdown
Contributor Author

Review round complete (for the morning review): the review found 2 must-fix bugs - foreach nodes with mixed instance outcomes never reached the failure policy, and stop() raced the control loop (run could finalize as done after an explicit stop). Both are fixed in commit 8344cdb with 7 regression tests (suite now 32, all green; gates clean): foreach now fails at the first permanently failed instance (fail_fast cancels in-flight siblings immediately), stop() uses a transitional stopping state with finalize guards, and resume() no longer sleeps in the model turn. Both fixes were verified by re-running the original reproductions. Remaining nits are documented V1 limitations in the PR body (settlement-only budgets, cancellation shortcut reporting).

@sethkarten

Copy link
Copy Markdown
Contributor Author

Integration note (verified by merging all open swarm branches together): main now enforces a deferred-boot contract - repl tests assert Version: ImageMagick 7.1.2-31 Q16-HDRI aarch64 8309dc92a:20260903 https://imagemagick.org
Copyright: (C) 1999 ImageMagick Studio LLC
License: https://imagemagick.org/license/
Features: Cipher DPC HDRI Modules
Delegates (built-in): bzlib freetype heic jng jpeg lcms ltdl lzma png tiff webp xml zlib zstd
Compiler: clang (21.0.0)
Usage: import [options ...] [ file ]

Image Settings:
-adjoin join images into a single multi-image file
-border include window border in the output image
-channel type apply option to select image channels
-colorspace type alternate image colorspace
-comment string annotate image with comment
-compress type type of pixel compression when writing the image
-define format:option
define one or more image format options
-density geometry horizontal and vertical density of the image
-depth value image depth
-descend obtain image by descending window hierarchy
-display server X server to contact
-dispose method layer disposal method
-dither method apply error diffusion to image
-delay value display the next image after pausing
-encipher filename convert plain pixels to cipher pixels
-endian type endianness (MSB or LSB) of the image
-encoding type text encoding type
-filter type use this filter when resizing an image
-format "string" output formatted image characteristics
-frame include window manager frame
-gravity direction which direction to gravitate towards
-identify identify the format and characteristics of the image
-interlace type None, Line, Plane, or Partition
-interpolate method pixel color interpolation method
-label string assign a label to an image
-limit type value Area, Disk, Map, or Memory resource limit
-monitor monitor progress
-page geometry size and location of an image canvas
-pause seconds seconds delay between snapshots
-pointsize value font point size
-quality value JPEG/MIFF/PNG compression level
-quiet suppress all warning messages
-regard-warnings pay attention to warning messages
-repage geometry size and location of an image canvas
-respect-parentheses settings remain in effect until parenthesis boundary
-sampling-factor geometry
horizontal and vertical sampling factor
-scene value image scene number
-screen select image from root window
-seed value seed a new sequence of pseudo-random numbers
-set property value set an image property
-silent operate silently, i.e. don't ring any bells
-snaps value number of screen snapshots
-support factor resize support: > 1.0 is blurry, < 1.0 is sharp
-synchronize synchronize image to storage device
-taint declare the image as modified
-transparent-color color
transparent color
-treedepth value color tree depth
-verbose print detailed information about the image
-virtual-pixel method
Constant, Edge, Mirror, or Tile
-window id select window with this id or name
root selects whole screen

Image Operators:
-annotate geometry text
annotate the image with text
-colors value preferred number of colors in the image
-crop geometry preferred size and location of the cropped image
-encipher filename convert plain pixels to cipher pixels
-extent geometry set the image size
-geometry geometry preferred size or location of the image
-help print program options
-monochrome transform image to black and white
-negate replace every pixel with its complementary color
-quantize colorspace reduce colors in this colorspace
-resize geometry resize the image
-rotate degrees apply Paeth rotation to the image
-strip strip image of all profiles and comments
-thumbnail geometry create a thumbnail of the image
-transparent color make this color transparent within the image
-trim trim image edges
-type type image type

Miscellaneous Options:
-debug events display copious debugging information
-help print program options
-list type print a list of supported option arguments
-log format format of debugging information
-version print version information

By default, 'file' is written in the MIFF image format. To
specify a particular image format, precede the filename with an image
format name and a colon (i.e. ps:image) or specify the image type as
the filename suffix (i.e. image.ps). Specify 'file' as '-' for
standard input or output. must not pull asyncio or secrets. prime-agent-runtime/src/rlm/swarm.py imports asyncio at module level; defer it into the four functions that use it (SwarmExecutor ctor default sleep, _start_loop, _control_loop, _loop_body) when refreshing onto main.

sethkarten added a commit that referenced this pull request Sep 18, 2026
…p race)

Address the PR #2401 review findings:

- foreach mixed-instance outcomes (finding 1): a node now fails at its
  FIRST permanently failed instance instead of waiting for all instances
  to settle. Previously a failure that settled before its siblings left
  the node stuck in running with every instance terminal and the failure
  policy never applied (the run ended via the stall guard); fail_fast
  could also never cancel in-flight siblings. The policy guard keeps the
  second and later permanent failures no-ops.
- stop() race (finding 2): stop() sets a transitional "stopping" state
  before its first await; the control loop re-checks state before the
  completion check, _finalize only runs while the run is running, the
  node-failure policy records node errors but never overwrites a state
  another transition owns, and the executor-error path no longer
  overwrites a concurrent stop with failed. A stopped run now stays
  stopped: no new children admitted during the stop window, no spurious
  done + finished notice. Repeated stop() is idempotent (finding 7) and
  appends only one run_stopped ledger event.
- resume() defers rate-limited admissions to the control loop with
  backoff exactly like run() (finding 3), so it never sleeps in the
  calling model turn.
- fail_fast milestone detail now counts cancelled in-flight children.
- tests: the host fake is now an async callable patched in directly
  (AsyncMock does not await async side effects) with per-child-id
  scripted outcomes and collect/delete suspension gates (finding 6) so
  races and mixed instances reach real paths; new tests cover mixed-
  instance foreach for every policy, the gated stop race, repeated stop,
  two-parent fan-in in one collect batch (finding 8), the
  ANSWER_CAPTURE_CAP slice (finding 9), and the resume deferral path.
@sethkarten
sethkarten force-pushed the swarm/swarm-dag-executor branch from 8344cdb to 39dac8d Compare September 18, 2026 06:09
@sethkarten

Copy link
Copy Markdown
Contributor Author

Consolidated into #2452 (DAG swarm workflows as a single review/merge unit) per the DAG-first review plan. All commits, tests, and review history from this PR are included there; this branch is unchanged and can be deleted.

@sethkarten sethkarten closed this Sep 18, 2026
@sethkarten sethkarten reopened this Sep 18, 2026
sethkarten added a commit that referenced this pull request Sep 18, 2026
…p race)

Address the PR #2401 review findings:

- foreach mixed-instance outcomes (finding 1): a node now fails at its
  FIRST permanently failed instance instead of waiting for all instances
  to settle. Previously a failure that settled before its siblings left
  the node stuck in running with every instance terminal and the failure
  policy never applied (the run ended via the stall guard); fail_fast
  could also never cancel in-flight siblings. The policy guard keeps the
  second and later permanent failures no-ops.
- stop() race (finding 2): stop() sets a transitional "stopping" state
  before its first await; the control loop re-checks state before the
  completion check, _finalize only runs while the run is running, the
  node-failure policy records node errors but never overwrites a state
  another transition owns, and the executor-error path no longer
  overwrites a concurrent stop with failed. A stopped run now stays
  stopped: no new children admitted during the stop window, no spurious
  done + finished notice. Repeated stop() is idempotent (finding 7) and
  appends only one run_stopped ledger event.
- resume() defers rate-limited admissions to the control loop with
  backoff exactly like run() (finding 3), so it never sleeps in the
  calling model turn.
- fail_fast milestone detail now counts cancelled in-flight children.
- tests: the host fake is now an async callable patched in directly
  (AsyncMock does not await async side effects) with per-child-id
  scripted outcomes and collect/delete suspension gates (finding 6) so
  races and mixed instances reach real paths; new tests cover mixed-
  instance foreach for every policy, the gated stop race, repeated stop,
  two-parent fan-in in one collect batch (finding 8), the
  ANSWER_CAPTURE_CAP slice (finding 9), and the resume deferral path.
@sethkarten
sethkarten force-pushed the swarm/swarm-dag-executor branch from 8f57a69 to b037304 Compare September 18, 2026 06:58
@sethkarten sethkarten changed the title swarm: swarm DAG executor (rlm.swarm run/status/stop, nonblocking control loop) swarm: state-machine swarm executor (guarded transitions, re-entry, wait states) Sep 18, 2026
sethkarten added a commit that referenced this pull request Sep 18, 2026
…p race)

Address the PR #2401 review findings:

- foreach mixed-instance outcomes (finding 1): a node now fails at its
  FIRST permanently failed instance instead of waiting for all instances
  to settle. Previously a failure that settled before its siblings left
  the node stuck in running with every instance terminal and the failure
  policy never applied (the run ended via the stall guard); fail_fast
  could also never cancel in-flight siblings. The policy guard keeps the
  second and later permanent failures no-ops.
- stop() race (finding 2): stop() sets a transitional "stopping" state
  before its first await; the control loop re-checks state before the
  completion check, _finalize only runs while the run is running, the
  node-failure policy records node errors but never overwrites a state
  another transition owns, and the executor-error path no longer
  overwrites a concurrent stop with failed. A stopped run now stays
  stopped: no new children admitted during the stop window, no spurious
  done + finished notice. Repeated stop() is idempotent (finding 7) and
  appends only one run_stopped ledger event.
- resume() defers rate-limited admissions to the control loop with
  backoff exactly like run() (finding 3), so it never sleeps in the
  calling model turn.
- fail_fast milestone detail now counts cancelled in-flight children.
- tests: the host fake is now an async callable patched in directly
  (AsyncMock does not await async side effects) with per-child-id
  scripted outcomes and collect/delete suspension gates (finding 6) so
  races and mixed instances reach real paths; new tests cover mixed-
  instance foreach for every policy, the gated stop race, repeated stop,
  two-parent fan-in in one collect batch (finding 8), the
  ANSWER_CAPTURE_CAP slice (finding 9), and the resume deferral path.
@sethkarten
sethkarten force-pushed the swarm/swarm-dag-executor branch from b037304 to 706fb73 Compare September 18, 2026 07:25
@sethkarten

Copy link
Copy Markdown
Contributor Author

Status correction: the consolidation above was reverted - this PR is the active review unit again, restored as part of the stacked series (this PR -> its dependents up to #2485). Ignore the "Consolidated into #2452" note; #2452 is closed and its branch deleted.

@sethkarten
sethkarten added this pull request to stack #2959 September 27, 2026 20:01
sethkarten added a commit that referenced this pull request Sep 27, 2026
…p race)

Address the PR #2401 review findings:

- foreach mixed-instance outcomes (finding 1): a node now fails at its
  FIRST permanently failed instance instead of waiting for all instances
  to settle. Previously a failure that settled before its siblings left
  the node stuck in running with every instance terminal and the failure
  policy never applied (the run ended via the stall guard); fail_fast
  could also never cancel in-flight siblings. The policy guard keeps the
  second and later permanent failures no-ops.
- stop() race (finding 2): stop() sets a transitional "stopping" state
  before its first await; the control loop re-checks state before the
  completion check, _finalize only runs while the run is running, the
  node-failure policy records node errors but never overwrites a state
  another transition owns, and the executor-error path no longer
  overwrites a concurrent stop with failed. A stopped run now stays
  stopped: no new children admitted during the stop window, no spurious
  done + finished notice. Repeated stop() is idempotent (finding 7) and
  appends only one run_stopped ledger event.
- resume() defers rate-limited admissions to the control loop with
  backoff exactly like run() (finding 3), so it never sleeps in the
  calling model turn.
- fail_fast milestone detail now counts cancelled in-flight children.
- tests: the host fake is now an async callable patched in directly
  (AsyncMock does not await async side effects) with per-child-id
  scripted outcomes and collect/delete suspension gates (finding 6) so
  races and mixed instances reach real paths; new tests cover mixed-
  instance foreach for every policy, the gated stop race, repeated stop,
  two-parent fan-in in one collect batch (finding 8), the
  ANSWER_CAPTURE_CAP slice (finding 9), and the resume deferral path.
@sethkarten
sethkarten force-pushed the swarm/swarm-dag-executor branch from 706fb73 to fb55a82 Compare September 27, 2026 20:09
Add the executor for stored swarm DAG entries on top of PR #2397's
validator (rlm/swarm.py) and expose it as the rlm.swarm namespace
(run/status/stop/resume) in rlm/__init__.py.

The SwarmExecutor runs a canonicalized DAG through the existing RLM
supervisor: nodes are admitted with rlm.spawn, settled through
rlm.collect, and cancelled with rlm.delete_subagent. The supervisor owns
the children; the executor owns run state in kernel memory. run() is a
nonblocking admission phase: it re-validates the stored dag, resolves
every subagent reference (reporting all failures before starting
anything), reports the resolved node count and max_parallel, starts
every ready node up to max_parallel, and returns while a background
asyncio task continues the run. The control loop is iterative (tested
at 100-node chain and 1000-node fan scale), polls collect with a 2s
timeout, binds inputs ({name} placeholder render plus an appended
Inputs section), expands foreach nodes with clamped instances, retries
child failures, enforces node and run wall-clock budgets via an
injectable clock, applies fail_fast/continue/escalate failure policies,
and backs off rate-limited spawn admissions (doubling, capped at 60s,
max 5 attempts) before failing the node.

Progress reaches the parent through one quiet notice per run milestone
(finished/failed/paused/budget-exceeded) via a new "swarm.progress" host
handler in agent-session.ts (steer lane, mirroring bash.completed), and
the run's event ledger records spawned/settled/answer_captured/retry/
error events with arrived/shown/delivered stages; status() returns node
states plus the trailing 50 events and marks the ledger delivered.

Runs do not survive a kernel restart (children are supervisor-owned and
keep running); captured answers are collect previews capped by the host
at 160 characters (compactRlmText).
…p race)

Address the PR #2401 review findings:

- foreach mixed-instance outcomes (finding 1): a node now fails at its
  FIRST permanently failed instance instead of waiting for all instances
  to settle. Previously a failure that settled before its siblings left
  the node stuck in running with every instance terminal and the failure
  policy never applied (the run ended via the stall guard); fail_fast
  could also never cancel in-flight siblings. The policy guard keeps the
  second and later permanent failures no-ops.
- stop() race (finding 2): stop() sets a transitional "stopping" state
  before its first await; the control loop re-checks state before the
  completion check, _finalize only runs while the run is running, the
  node-failure policy records node errors but never overwrites a state
  another transition owns, and the executor-error path no longer
  overwrites a concurrent stop with failed. A stopped run now stays
  stopped: no new children admitted during the stop window, no spurious
  done + finished notice. Repeated stop() is idempotent (finding 7) and
  appends only one run_stopped ledger event.
- resume() defers rate-limited admissions to the control loop with
  backoff exactly like run() (finding 3), so it never sleeps in the
  calling model turn.
- fail_fast milestone detail now counts cancelled in-flight children.
- tests: the host fake is now an async callable patched in directly
  (AsyncMock does not await async side effects) with per-child-id
  scripted outcomes and collect/delete suspension gates (finding 6) so
  races and mixed instances reach real paths; new tests cover mixed-
  instance foreach for every policy, the gated stop race, repeated stop,
  two-parent fan-in in one collect batch (finding 8), the
  ANSWER_CAPTURE_CAP slice (finding 9), and the resume deferral path.
…, re-entry, wait states)

run() canonicalizes the stored spec to machine form (dag sugar compiles),
enters the entry states, and drives the machine to quiescence: every settle
is evaluated once and ALL guard-passing transitions fire (fan-out legal);
max_entries blocking is recorded as transition_blocked events; self-loops
and back-edges re-enter states with freshly re-bound inputs; foreach expands
per entry; budgets/retries/failure policies stay V1 (per instance, per
entry); wait states register rlm.watch.path/agent watches and settle the
implicit event output (detail, timed_out) through the injectable clock;
max_transitions pauses once with a budget_exceeded milestone. status() adds
per-state entries_used/max_entries, per-entry breakdowns, and a
transitions_fired count. Every V1 dag executor behavior is preserved
through the compiler path: all 32 existing executor tests run unchanged.
Cover the state-machine semantics end to end with the fake host: a
three-round review loop (guarded switch approved=false twice then true;
3 reviewing entries, 2 fixing entries, done), a mutually exclusive guard
switch where only the matching branch enters, fan-out from one settle,
max_entries blocking recorded as transition_blocked events with a
quiescent done, a self-loop re-entering until its guard fails, wait-state
settles from a completed path watch and from the injectable-clock timeout
(event object bound into the dependent prompt), a max_transitions pause
reported once and resumed, and stop() cancelling an active watch.
…ume, JSON-strict guards)

Wait states are gated at the validator (the rlm.watch.* host handlers
arrive with the communication series #2351/#2356), so the executor's wait
registration/timeout/settle/cancel paths and their tests are removed
entirely; quiescence, re-entry, and fan-out are unchanged. Review fixes:
a max_transitions pause mid-settle now records the transition index and a
resume continues after it (previously the settle re-fired already-fired
transitions, duplicating spawned work); eq/ne guards compare JSON-strictly
(a bool never equals a number, numbers compare numerically); contains
requires a non-empty list value at validation and is defensively false at
runtime; the max_transitions pause fires its own max_transitions_exceeded
milestone kind (no collision with the run-budget budget_exceeded, both
notice-able); the control-loop stall reports the pending entry whose input
source never settled, with a test; EVENT_WINDOW grows 50 -> 200 (the
pr-manager happy path is ~43 events before any retry); refinement
validateEdit accepts machine-form swarm edits (either dag or machine
object, never both). Also adds optional inputs: an input flagged optional
binds a null sentinel instead of waiting when its source state never
settled, which lets loop states (review with the previous fix report) run
their first round before the fixer exists; compiled dags never emit it.
@sethkarten
sethkarten force-pushed the swarm/swarm-dag-executor branch from fb55a82 to 2c6e883 Compare September 27, 2026 21:24
@sethkarten sethkarten changed the title swarm: state-machine swarm executor (guarded transitions, re-entry, wait states) factory: state-machine workflow executor (factory.run/status/stop/resume, guarded transitions, re-entry) Sep 27, 2026
…ests

check:test-policy flags every new sleep()/delay() call in test files, and
asyncio.sleep(0) in the new factory executor tests tripped it on all three
stack branches. Replace those cooperative yields with a yield_loop_turn()
helper (call_soon + Event) that yields the loop with no wall-clock wait and
no flagged token.
…observable stages

Reviewer batch 1 findings on the factory executor:

F1 (high): a child admitted after stop() was never cancelled. _admit now
re-checks the run state after every await: a spawn that lands in a
non-running run is retracted (delete_subagent + cancelled instance,
never registered running) and a backoff that wakes in a non-running run
never retries; both return stopped and _spawn_ready stops admitting.
Pinned by test_stop_during_in_flight_admission_deletes_the_child and
test_stop_during_admission_backoff_cancels_the_instance (FakeHost can
now gate rlm.run admissions; GatedSleep suspends the backoff window).

F2 (medium): _milestone returned before _event, so a repeated milestone
(escalate pause -> resume -> fail again) left no ledger record. Every
milestone is now appended; only the parent notice stays deduped per
kind. Pinned by test_repeated_pause_records_every_milestone_in_the_ledger.

F3 (question): documented that _NodeInstance.attempt counts admission
attempts (429 deferrals included) and retries compares against it.

F4 (question): _parse_json_output tried only the trailing fence; it now
scans the fences trailing-first for the one carrying the output port,
then falls back to the whole answer. Pinned by
test_json_output_binds_from_the_fence_that_carries_the_port.

F5 (question): status() collapsed every stage to delivered; it now
advances only recorded/arrived events, so shown stays observable and
the stage taxonomy is readable through status().
_halt_nonterminal cancelled the entry but left its prepared instances
status pending, so a stopped/failed run reported pending instances the
replay checker (factory-eval checkReplayLedger) rejects as impossible.
The instance is now cancelled with its own ledger event (no child: it
was cancelled before admission); _admit's stop-window retraction covers
the in-flight case and deletes the child it returned.
…ations after failed deletes

Reviewer round-2 findings on the factory executor:

N1 (low): _halt_nonterminal built its cancellation list from state.status,
which reports the LATEST entry only, so a re-entered state whose earlier
entry was still in flight (later entry settled) was never reported or
cancelled: stop() returned cancelled: [] and the stopped ledger still
showed a running entry. The list is now built from entries (any
non-terminal entry) plus never-entered states, so stop()/fail_fast
cancel and report every in-flight entry. Pinned by
test_stop_cancels_an_in_flight_earlier_entry_of_a_reentered_state
(stop() -> cancelled ['x'], entries read [(0, cancelled), (1, done)]).

N2 (low): a delete_subagent failure recorded only cancel_failed while
the instance still reads cancelled, so the eval replay checker rejected
the ledger (cancelled without a cancel event; spawned never cancelled).
The ledger now records the cancellation alongside the failure (the slot
is released either way), in both _cancel_running and _retract_admission.
Pinned by test_failed_delete_records_cancel_failed_and_cancelled.
@sethkarten
sethkarten marked this pull request as ready for review September 28, 2026 02:41
"(the original DAG sugar in arguments['dag'] compiles to machine form): manage them with "
"create_factory/update_factory/delete_factory (create_factory validates either form at write time); run "
"them with rlm.factory.run(\"<id>\") once the executor lands in a follow-up PR.",
"them with await rlm.factory.run(\"<id>\"), watch with rlm.factory.status(run_id), stop with "

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium rlm/harness.py:1054

The overview instructs callers to invoke rlm.factory.status, rlm.factory.stop, and rlm.factory.resume without await, so the REPL returns unexecuted coroutine objects instead of observing, stopping, or resuming the run. Add await to each async namespace call; otherwise a resident factory run can continue consuming children.

Also found in 2 other location(s)

packages/coding-agent/src/core/prompts/rlm.ts:49

The new prompt omits await for rlm.factory.status(run_id), rlm.factory.stop(run_id), and rlm.factory.resume(run_id), although all three namespace methods are async. Following this instruction in the Python REPL merely creates unexecuted coroutine objects, so the model cannot actually inspect, stop, or resume a factory run.

packages/coding-agent/src/core/refinement/refinement.ts:721

The newly added factory usage hint omits await for both rlm.factory.status(run_id) and rlm.factory.stop(run_id), although those methods are declared async in the runtime. A model following this prompt merely creates unexecuted coroutine objects, so it neither observes a run nor stops it; in particular, a resident or stuck factory run continues consuming children.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @prime-agent-runtime/src/rlm/harness.py around line 1054:

The overview instructs callers to invoke `rlm.factory.status`, `rlm.factory.stop`, and `rlm.factory.resume` without `await`, so the REPL returns unexecuted coroutine objects instead of observing, stopping, or resuming the run. Add `await` to each async namespace call; otherwise a resident factory run can continue consuming children.

Evidence trail:
Reviewed commit c544820f. prime-agent-runtime/src/rlm/harness.py:1041-1055; prime-agent-runtime/src/rlm/__init__.py:531-551; prime-agent-runtime/src/rlm/repl.py:552-574; prime-agent-runtime/src/rlm/factory.py:1191-1206.

Also found in 2 other location(s):
- packages/coding-agent/src/core/prompts/rlm.ts:49 -- The new prompt omits `await` for `rlm.factory.status(run_id)`, `rlm.factory.stop(run_id)`, and `rlm.factory.resume(run_id)`, although all three namespace methods are async. Following this instruction in the Python REPL merely creates unexecuted coroutine objects, so the model cannot actually inspect, stop, or resume a factory run.
- packages/coding-agent/src/core/refinement/refinement.ts:721 -- The newly added factory usage hint omits `await` for both `rlm.factory.status(run_id)` and `rlm.factory.stop(run_id)`, although those methods are declared `async` in the runtime. A model following this prompt merely creates unexecuted coroutine objects, so it neither observes a run nor stops it; in particular, a resident or stuck factory run continues consuming children.

Comment on lines +1976 to +1977
if state.lifecycle == "resident" and entry.status == "running":
continue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High rlm/factory.py:1976

_run_complete reports quiescence while a resident running entry still has pending instances, so the run finalizes and exits before those children are admitted. This occurs when max_parallel leaves a resident transition queued; exclude resident entries with pending instances from the quiescent case.

Suggested change
if state.lifecycle == "resident" and entry.status == "running":
continue
if state.lifecycle == "resident" and entry.status == "running" and not any(
instance.status == "pending" for instance in entry.instances
):
continue
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @prime-agent-runtime/src/rlm/factory.py around lines 1976-1977:

`_run_complete` reports quiescence while a resident `running` entry still has `pending` instances, so the run finalizes and exits before those children are admitted. This occurs when `max_parallel` leaves a resident transition queued; exclude resident entries with pending instances from the quiescent case.

Evidence trail:
prime-agent-runtime/src/rlm/factory.py:1465-1490, 1508-1512, 1966-1979, 2068-2073 at commit 3daf330

Comment on lines +1592 to +1594
for entry in state.entries:
for instance in entry.instances:
if instance.status == "pending":

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High rlm/factory.py:1592

_next_pending_instance admits queued siblings even after a foreach entry has become terminal error, so failure_policy: "continue" still runs additional child work after the node failed. Skip pending instances belonging to non-running entries.

             for entry in state.entries:
+                if entry.status != "running":
+                    continue
                 for instance in entry.instances:
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @prime-agent-runtime/src/rlm/factory.py around lines 1592-1594:

`_next_pending_instance` admits queued siblings even after a `foreach` entry has become terminal `error`, so `failure_policy: "continue"` still runs additional child work after the node failed. Skip pending instances belonging to non-running entries.

Evidence trail:
c544820
prime-agent-runtime/src/rlm/factory.py:1465-1489
prime-agent-runtime/src/rlm/factory.py:1589-1596
prime-agent-runtime/src/rlm/factory.py:1785-1858
prime-agent-runtime/src/rlm/factory.py:1966-1979

if self._run_complete(run):
await self._finalize(run)
return
if run.run_budget_ms is not None and not run.budget_reported:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High rlm/factory.py:2074

The initial admission phase can launch instances after run_budget_ms has expired, so a slow spawn fills max_parallel instead of stopping new work at the budget boundary. The budget check at line 2074 only runs in the background control loop, after run() has already called _spawn_ready; enforce the budget before each admission (including that initial path).

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @prime-agent-runtime/src/rlm/factory.py around line 2074:

The initial admission phase can launch instances after `run_budget_ms` has expired, so a slow `spawn` fills `max_parallel` instead of stopping new work at the budget boundary. The budget check at line 2074 only runs in the background control loop, after `run()` has already called `_spawn_ready`; enforce the budget before each admission (including that initial path).

Evidence trail:
prime-agent-runtime/src/rlm/factory.py:1108-1112, 1304-1308, 1471-1481, 1624-1628, 2074-2089 at commit c544820f. Repository: https://github.com/PrimeIntellect-ai/prime-agent. Verification command: git show c544820f:prime-agent-runtime/src/rlm/factory.py

f"resume with await rlm.factory.resume('{run.run_id}')",
)
return
started = await self._spawn_ready(run, allow_backoff=True)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High rlm/factory.py:2089

_spawn_ready(..., allow_backoff=True) blocks the sole control loop during its rate-limit retry sleeps, so already-running children are not collected and run-budget checks are delayed for the entire backoff (up to 15 seconds). A child that completes during this interval is therefore observed late and can be classified as exceeding its per-state budget. Make admission backoff non-blocking to the control loop, or move retries into a separate task so collection and budget processing continue.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @prime-agent-runtime/src/rlm/factory.py around line 2089:

`_spawn_ready(..., allow_backoff=True)` blocks the sole control loop during its rate-limit retry sleeps, so already-running children are not collected and run-budget checks are delayed for the entire backoff (up to 15 seconds). A child that completes during this interval is therefore observed late and can be classified as exceeding its per-state budget. Make admission backoff non-blocking to the control loop, or move retries into a separate task so collection and budget processing continue.

Evidence trail:
prime-agent-runtime/src/rlm/factory.py:2046-2089, 1621-1643, 764-771, 1689-1724 at commit c544820f

Comment on lines +1200 to +1201
if run.state == "stopped":
return {"run_id": run.run_id, "state": "stopped", "cancelled": []}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium rlm/factory.py:1200

Concurrent stop() calls both enter _halt_nonterminal while the first call is awaiting child deletion, so the same children receive duplicate delete_subagent requests and a second run_stopped event is recorded. Guard the transitional "stopping" state as well as "stopped" before starting another cancellation pass.

Suggested change
if run.state == "stopped":
return {"run_id": run.run_id, "state": "stopped", "cancelled": []}
if run.state in ("stopping", "stopped"):
return {"run_id": run.run_id, "state": run.state, "cancelled": []}
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @prime-agent-runtime/src/rlm/factory.py around lines 1200-1201:

Concurrent `stop()` calls both enter `_halt_nonterminal` while the first call is awaiting child deletion, so the same children receive duplicate `delete_subagent` requests and a second `run_stopped` event is recorded. Guard the transitional `"stopping"` state as well as `"stopped"` before starting another cancellation pass.

Evidence trail:
Reviewed commit c544820. prime-agent-runtime/src/rlm/factory.py:1191-1206, 1868-1909, 1939-1962; prime-agent-runtime/src/rlm/__init__.py:460-479; prime-agent-runtime/test/test_factory_executor.py:964-968. Verify with `git show c544820 -- prime-agent-runtime/src/rlm/factory.py`.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread prime-agent-runtime/src/rlm/factory.py
Comment thread prime-agent-runtime/src/rlm/factory.py

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread prime-agent-runtime/src/rlm/factory.py

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 6dfd619. Configure here.

parts.append(f"i{instance_index}")
if attempt > 1:
parts.append(f"a{attempt}")
return "-".join(parts)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Truncated child names can collide

Medium Severity

_child_name keeps only the first 20 characters of the state id, so two valid slugs that share a prefix (for example collect-findings-pass-1 and collect-findings-pass-2) produce the same sibling name. rlm.spawn requires unique sibling names, so the second admission fails and the state's failure_policy applies.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 6dfd619. Configure here.

for instance in entry.instances:
if instance.status == "pending":
return state, entry, instance
return None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Failed foreach still admits siblings

Medium Severity

_next_pending_instance returns any pending instance and does not check whether its entry already failed. After a foreach entry hits a permanent failure under continue or escalate, later items that were never admitted are still spawned. Those children consume max_parallel slots and do extra work for an entry that is already terminal.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 6dfd619. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant