Skip to content

Phase 3: every SSH host runs a managed Orca server (orcad), replacing the relay, and orca serve runs on it - #24863

Merged
OrcaWin merged 188 commits into
mainfrom
OrcaWin/node-rt-phase3
Oct 7, 2026
Merged

OrcaWin merged 188 commits into
mainfrom
OrcaWin/node-rt-phase3

Conversation

@OrcaWin

@OrcaWin OrcaWin commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator
Files Added Deleted Net
Test 367 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​36316 $\color{#cf222e}{\Huge{\mathbf{−}}}$​1109 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​35207
Prod 597 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​39851 $\color{#cf222e}{\Huge{\mathbf{−}}}$​3649 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​36202

ELI5

Today Orca installs a small helper (the relay) on every SSH host, and that install often breaks on hosts with the wrong Node version or no build tools. After this PR, every SSH host runs Orca's own server program (orcad) instead. The app ships orcad with a pinned Node and sets it up when you connect.

Hosts you already use move over on their next connect. Orca asks first if that would restart open terminals. If a host can't run orcad, Orca keeps using the relay and says why. There is no setting to turn any of this on. orca serve runs on orcad too.

This was built and reviewed as about 135 smaller PRs on an integration branch. It lands as one change so main never carries half of it.

What Changed

Connecting to an SSH host

Before: every connect installed or reused the relay.

After: before any relay starts, the connect picks the host's server.

Host What happens
Empty Sets up orcad and connects to it.
Has Orca projects, no running terminals Moves its projects, folders, project groups, open editor tabs and their scrollback onto orcad, then connects.
Already moved Connects straight to orcad.
Has open terminals from this desktop Stays on the relay, with a once-per-version toast and a status-line button: "Move to a managed Orca server… its N open terminals will restart." Confirming stops those terminals, proves they exited, and reconnects to do the move. If the move is refused, the host reconnects on its relay.
Has open terminals only from another Orca desktop or session Stays on the relay and says whose terminals they are. It offers no Move, so this desktop never restarts another desktop's shells.
Can't run orcad Stays on the relay, and SSH Hosts says why.

Reasons a host can't run orcad: unsupported platform, no orcad build in this app, failed native or C-library check, failed runtime self-test, or Windows security software. Orca remembers these per app version. Every other failure retries on the next connect, and a failed attempt always releases its lock.

How the move works: it's journaled and crash-safe: lock the host, set up the server, copy, commit once. An interrupted move finishes or backs out cleanly on the next connect. Each host's saved state stays its own, even when two hosts share a repo id.

Counting terminals honestly: a running shell can't be carried into another server process, so the move needs proof that every relay terminal has exited.

  • Orca asks every relay on the host. On Windows that includes every relay pipe, so a terminal another computer opened also counts. A relay shell that Orca has no record of counts as live.
  • A terminal Orca can't check counts as running, never as gone. A missing answer from the relay never proves an exit.
  • Once the relay session can answer, the connect checks again. Records of terminals the relays prove ended are cleared, so the next connect converts.

How the app reaches orcad: through an SSH port forward to the port orcad really bound, after checking that whatever answers is this host's orcad, by its encryption key and our login token. orcad gets a new internal id on every restart, so a desktop reconnects fine after another desktop restarted or updated the server. Where the SSH server forbids port forwarding, the app uses a small helper on the normal SSH connection instead (a "stdio bridge").

Old Linux (CentOS 7, glibc 2.17): fully managed. It gets a runtime and file watcher built for old glibc, and its relay fallback uses Orca's own Node, not the host's.

Keeping a managed server running

  • Updates: on every connect, and when the app restores its servers at launch, a server on an older build is updated to the app's build.
    • The update waits until no terminals are running, or can be counted.
    • A failed update keeps the old build serving, and it isn't retried until the app version changes.
    • A server set up by a newer Orca is never downgraded, and a build you rolled back from is never reapplied.
    • A rollback restores the older snapshot only once orcad confirms no terminals are running. An update that would end a degraded host's terminals is refused.
    • One owner at a time: every update, rollback, stop, recover and restart from idle holds the host's lock with its own owner token. Each step re-checks the token, and the lock is released only by its owner. So a desktop that was suspended mid-update can't clobber another desktop's update when it wakes. This also fixes BUG-21: a leftover lock with nothing interrupted behind it is no longer read as "Recover it first".
    • Snapshots: saving, restoring or clearing a server's saved-data snapshot waits as long as it needs, runs one at a time on the host, and is owned by its process group. Only proof that it exited frees the lock, so a slow restore can't be raced by a second one.
  • Idle shutdown and restart:
    • An orcad that Orca launched exits after 15 minutes with no clients, terminals, working agents, scheduled automations, or migrations and updates in progress. Terminals are never killed.
    • When orcad isn't answering, whatever stopped it (idle, a crash, a reboot), the app restarts it on the next connect, call or wake. It does this only once the host proves the old process exited.
    • If the connection drops just as a restart takes the host's lock, the reconnected restart clears its own lock instead of waiting for it to go stale. The lock's owner is recorded in the same command that creates it.
  • Automations run on orcad, both on managed hosts and under orca serve. A managed host closes finished automation terminals that nobody opened, keeping the newest 3 per automation. Before an update counts running terminals, it closes those finished, unused automation shells, so they don't block the update.

Terminals from the previous Orca version

Before: after an app update, a terminal the old relay still ran was replaced by an empty shell. The real shell kept running out of sight.

After:

  • The tab reattaches through the old relay's own connector, so scrollback, typing and output work.
  • Closing it stops it on the host and confirms the exit.
  • Input that no connection can serve is refused, not silently dropped.

Settings and CLI

  • SSH Hosts:
    • A status line only when something needs attention (setup progress, an update note, the relay and why, a failure), plus a Details section with the failure reason and the last lines of orcad's log. A healthy managed host shows no line (fix(ssh): keep the SSH host card quiet while its managed server is healthy #26072).
    • One row per machine: an SSH host and its managed server show as one host in the new-workspace picker, sidebar host sections and filter, Add Project, the jump palette, notification toggles and repository host setups (fix(hosts): show an SSH host and its managed Orca server as one host #26076). Both ids stay resolvable, so saved selections, folders and projects under either id keep working.
    • A managed host's connection settings can be edited.
    • Removing a managed host is refused before any terminal ends.
    • A stopped server disappears from its host without a reconnect, and leaves no empty project group.
  • Settings → Managed servers: always visible. It reads "unknown", not "Not running", when status never loaded.
  • CLI: new orca environment status | update | rollback | recover | stop | cancel-stop commands run the same actions from scripts. recover --accept-changed-state --yes restores the pre-update snapshot over changed data.
  • Removed: the experimental setting and the old manual move dialog.

Going back to an older Orca

  • The marker that says "this host is managed" lives in a new field that older builds keep but ignore. So an older build still lists the host and its projects, and uses them over the relay.
  • After a move, the host's old project records stay in the local profile. This build hides them, and older builds read them. Nothing deletes them automatically. They go away only when you stop the managed server (Stop… under Managed servers), remove the host, or uninstall.
  • On upgrading again, a host whose projects, folders and groups (by id and path) are unchanged stays on orcad.
    • A host where the older build added or removed any of those stays on the relay, marked "changed on an older Orca".
    • It offers two actions: Move the new projects, which moves only what was added (any conflict fails the whole move before commit), or Keep the server's version.
    • Nothing is merged automatically.
  • Edits an older build makes inside an already-moved project, such as a draft, don't mark the host as changed. They aren't lost: they stay in the retained records, which an older build, or Stop…, gives back.

orca serve

  • Runs on orcad by default on Linux, Windows and unpackaged macOS, and shares the desktop profile. ORCA_SERVE_RUNTIME=electron opts out, and every fallback prints its reason.
  • Packaged macOS stays on the desktop app's server, because only the desktop app can install updates to a machine that's serving.
  • On Windows, Orca resolves its AppData folders itself at startup. So it no longer crashes when Windows can't report roaming AppData, for example for orca serve started over SSH.

Security

The file orcad writes on a host to say it's ready holds the login token clients connect with. Under a common login umask, other users on a shared host could read it. Now it, the PID file and orcad's log are owner-only, and so are the folders that hold them. A sweep of other files Orca writes on hosts found no other readable secrets.

Disk use

  • On your computer: Orca keeps the orcad copy in use plus the two most recent others per platform. They stay across uninstall with the rest of the user data. A copy that's in use, being written, or serving a surviving terminal daemon is never removed.
  • On SSH hosts: after each setup or update, old orcad versions are removed only when proven stopped. If Orca can't prove that, it keeps them.

Remote compatibility

Clients and hosts update on their own schedules, so mixed versions are normal.

  • New server calls are used only when the other side advertises them. "Method not found" counts as a refusal, not an error.
  • Everything added to existing messages is an optional field, which older clients drop. The new Phase 3 messages accept fields from newer peers.
  • The terminal daemon's protocol version is unchanged. Two relay calls that never shipped were removed before they could.

Telemetry

Each connect sends one anonymous event: which server the host got, how the app reached it, why, and the host's OS, CPU and C library. Conversions, setup failures and the move prompt each get one event too. Every field comes from a fixed list, so no hostnames, paths or log text are sent.

Why

  • Why replace the relay: relay installs fail on hosts with the wrong Node or no build tools. One server package on a pinned Node removes that whole class of failures, and gives SSH and paired clients the same server.
  • Why decide on connect, with no setting: every host converges without user action, and a host that can't run orcad still works through the relay. The decision has to come before the relay starts, because the move can't lock a host that has a live relay session.
  • Why never move running terminals silently: a live shell can't move into a different server process. Killing it without asking would lose work. So Orca asks per host, and treats any terminal it can't verify as running.
  • Why keep old records, and never delete them automatically: hiding moved hosts from older builds would make a downgrade lose every converted host until the user upgraded again. An earlier design deleted the old records after a rollout flag flipped. Review after review found cases where that deletion could remove something a user had just written, so it was taken out (net 1,795 lines removed). Keeping the records costs some duplicate profile data, and possibly a relay running next to orcad on a downgraded machine. Safe automatic cleanup can come later, as its own change.
  • Why update only idle servers, and stop when idle: an update restarts the server, so it waits for a moment with no terminals running. Idle shutdown matches what the relay already did. Restarting on demand makes a stopped server harmless.
  • Why one PR: a partly wired main would convert hosts without the downgrade safety, the fallbacks or the recovery.

Linked Issue

Phase 3 of the Node runtime migration. Ports #16741, except Bun, live terminal handover, and moving hosts back off orcad.

Visual Proof

The new UI is one status line per host in Settings → SSH Hosts. Examples:

  • "Runs a managed Orca server"
  • "Runs the relay until its 1 open terminal is closed, then moves to a managed server", with a Move to managed server button
  • "Not moved to a managed server: "
  • "Changed on an older Orca…"

The UI also includes the move toast and dialog. Screenshots of the Managed servers section are in #24590.

Testing

  • I manually tested these changes locally
  • Automated tests added/updated, or explained why not below

Every constituent PR ran full CI, and ad hoc builds pin every job to one commit. The real-host testing drove the app and the orca CLI from ad hoc builds of this branch.

Linux and macOS real hosts

Thirteen rounds of testing ran on these hosts:

  • Debian amd64 and arm64
  • Alpine (musl)
  • CentOS 7
  • a host that forbids port forwarding
  • macOS localhost, with the desktop app running
  • relay-era hosts set up with v1.4.218, the last relay-only release
Scenario Result
Empty host becomes managed PASS on Debian, Alpine, CentOS 7 and macOS, in 20–96 s
Realistic v1.4.218 profile converts (repo, folder, group, editor tab) PASS. The tab loads, and the sidebar is clean.
Host with open terminals: offer, then Move PASS. Three relay shells, including one Orca had no record of, were stopped, and the host was managed in 30 s.
Previous-version terminal resumes PASS. Scrollback and typing work, on the same process.
Closing a reattached relay terminal PASS (ok, shell gone)
orcad killed, or host rebooted PASS. It comes back on its own in about 20 s, and terminals are kept.
Network drop PASS. It recovers on its own (3 of 3), and worktree ps shows "unverifiable" meanwhile.
No port forwarding PASS. Managed over the stdio bridge.
Downgrade to v1.4.218, then re-upgrade PASS. The host stays usable, and "Move the new projects" moves the additions.
Editing a managed host's port PASS. Same server, and the terminal is kept.
Update on connect PASS. Deferred when busy, updated when idle. --force works, and rollback with terminals open is refused.
Automation on orcad PASS. The run is dispatched on the host. A missing agent fails the run with the host's error instead of reading "completed".
Workspace port detection on a managed host (Debian and Alpine) PASS. A listener started in a managed terminal shows up under the right worktree, on its sidebar row and in the ports popover.
Update right after an idle restart, on a heavily loaded host PASS. Restarted, then updated to managed in 20 s, with rollback available.
Recover a host an interrupted update left wedged PASS. recover --accept-changed-state --yes restores it, and calls work.
A host whose only live terminal belongs to another desktop PASS. It isn't auto-converted.
orcad's file permissions on the host PASS. Its folder is owner-only (700), and the readiness file, PID file and log are owner-only (600).
orca environment stop --yes PASS. The server is removed, and the host clears without a reconnect.

Of the first 16 real-host bugs (BUG-1 to BUG-16), 15 are verified fixed in a rerun. BUG-12, a dropped network leaving a host stuck, didn't reproduce in 3 tries after its fix. The verified fixes include:

  • conversion refused by a selected workspace;
  • old terminals replaced by empty shells;
  • orcad never restarting;
  • moved tabs not loading;
  • CentOS 7 not connecting;
  • a port clash with another Orca;
  • a failed Move wedging the host;
  • a stale pairing after restart.

Later reruns found more:

The final real-host run, on the last ad hoc build before landing, passed:

Scenario Result
BUG-21: recovered host wakes and updates PASS
Two desktops racing one update PASS. One updates, and the other waits and backs off. Both end up on the managed server, with no lock left behind. The desktop that waited showed a stale "holds this host" status until it reconnected. It now clears that status once the host answers (#25995); real-host recheck: cleared about 55 s after the other desktop finished, where it used to stay 2–6+ min.
Automation terminal cleanup and the update gate PASS. Finished runs, failed ones included, are trimmed to the newest 3 per automation. A shell someone typed in is kept and still blocks a rollback. Unused shells are released so a rollback can go ahead.
Smoke: empty host managed, relay-era host converted on launch, close --all PASS

Other fixes that landed late:

Release upgrade and rollback (latest release v1.4.222 ↔ this branch)

Step Result
Release: connect, 2 long-running terminals PASS. Connected through the relay in 5 s; both shells kept running after quit.
Upgrade to this branch PASS. The host shows its 2 open relay terminals and offers Move; both reattached with output unbroken; Move took 7 s.
Roll back to the release PASS. It reconnects to its relay in 3 s and opens new terminals; the profile is intact. Terminals left on the managed server show as reconnecting (the expected one-way conversion).
Upgrade again PASS. Managed in 4 s; the managed server's terminals reattach live.

Windows real hosts

These ran as CI host cells: inbox and preview OpenSSH, on x64 and arm64, driven with the orca CLI.

Scenario Result
Empty host managed; relay host converts PASS on all four runners
Open relay terminal counted as live, with the offer shown PASS on all four (Windows bug 4 fixed)
Strict close of a reattached relay terminal, then convert PASS on all four, in two separate runs (ok in 438–654 ms, then the host converts)
orcad killed: automatic restart, terminal survives, reconnect stays managed PASS on all four
orca environment rm refused; orca environment stop --yes decommissions; reconnect redeploys PASS on all four
Workspace port detection PASS on all four. A listener on the host is detected with its process id, and matched to its workspace when the worktree path is in its command line.

All four Windows bugs found in this testing are fixed:

  • orcad never restarting;
  • the port-scan helper missing from orcad: on all four runners, the managed server now detects a listener on the host and matches it to its workspace.
  • a reattached terminal's close reading "unverifiable";
  • an open relay terminal counted as unverifiable instead of live.

Other evidence

  • CI conversion cells: a Linux Docker host and a Windows OpenSSH host convert from a profile shaped like the current release.
  • Downgrade check: v1.4.218's own code still lists a converted host and its project.
  • Switching orca serve between the desktop app and orcad: on Linux, macOS and Windows CI, a terminal survives both directions on one profile, and the right side holds the profile lock each time. On an installed Windows app, the relocated terminal daemon survives desktop → orcad → desktop with the same process.

Review loop

Every round reviewed the whole branch against main. Each finding was checked independently before it counted, and every confirmed finding was fixed in its own PR.

  • Rounds 1–8: reported every bug and every cleanup.
  • From round 9: each round runs two reviewers in parallel, working independently, and reports only P0 and P1 issues. Those are issues that would lose a user's work, leak data, leave a host unusable, or break a core flow. The loop stops when a round finds none on the branch.
Round Confirmed issues
1: five area reviews (SSH, orcad, migration, relay and CLI, UI and CI) Several dozen bugs, plus an estimated 3,000 lines of unused or duplicated code
2 47
3 11
4 8
5 8
6 7
7 5
Delta review of one batch of fixes 5
8 3 bugs, plus 4 cleanups
9 (P0/P1 only) 3 P1
10 (P0/P1 only) 0 on the branch, from both reviewers. 1 P1 was in a fix PR not yet merged at the time (#25811).
11 (P0/P1 only, final fix PRs) 1 P1, in #25811
12 (P0/P1 only) 3 P1, in #25811 and #25831. #25834 and #25825 clean.
13 (P0/P1 only) #25811 clean. 1 P1 each in #25834 and #25831.
Final targeted checks #25831 and #25834 clean
Second, independent reviewer, passes 1–10 (passes 6–10 P0/P1 only) 5, 3, 4, 4, 2, 1 P1, 0, 2 P1, 2 P1 (plus 1 already fixed), 0

The round 9 and pass 6 P1s were a snapshot restore racing a second restore, the readiness file's permissions, and the two-desktop restart identity. The readiness permissions and two-desktop identity are fixed (#25809, #25800). The restore race is fixed in #25811.

Where the loop ended: in round 10, both reviewers found the branch clean. Every fix made after that is clean on its final check.

The review fixes and cleanups removed several thousand lines. The largest single cut is the removal of automatic source retirement (net 1,795 lines).

Last fixes (rounds on each, Opus + Astra, P0/P1 only, all clean at their final heads): #26072 (SSH card status line), #26076 (one host row per machine; 4 P1s fixed, then clean), #26077 (Move keeps tabs; 3 P1s fixed, then clean), #26087 (Windows lock reclaim plus an instance-tied steal on every platform; 4 rounds, then clean).

Main sync: main was last merged into the branch at e547664 (#26108, #26147).

Known accepted limits

  • System ssh hosts that forbid port forwarding: a host reached only through the system ssh binary doesn't get the stdio bridge, so it stays on the relay.
  • Retained records are never cleaned up automatically: they stay in the local profile until Stop…, host removal or uninstall. Safe automatic cleanup is a later follow-up.
  • Windows port attribution: Windows can't read another process's working directory, so a port is matched to a workspace only when the worktree path is in the process's command line. A dev server started as node <repo>\… is matched; a bare listener shows as external. Local Windows works the same way today.
  • Upgrading from earlier test builds: a host that already ran orcad from an earlier build and has failed automation runs may need one orca environment update --force (or orca terminal close --all). The older server's cleanup rule doesn't release those shells. Hosts first set up by this build aren't affected.
  • A quit mid-update: on Linux, macOS and Windows hosts the next launch takes this desktop's own leftover locks back about 3 minutes after the old holder is proven exited (fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait #25941, fix(orcad): an install lock this desktop's exited process left mid-upload is taken over without the 20-minute wait #25991, fix(ssh): reclaim this desktop's own exited lock on Windows hosts too #26087). The takeover goes through the existing steal, holds the state-mutation lock while it replaces the fence, and every state change rechecks its fence after locking. Locks another desktop holds still wait out the 20-minute window. A state change from a pre-Phase-3 client can still be admitted late, a gap that predates this PR.
  • Unprovable lock holder: if a host can't prove who holds its lock, it stays busy until that holder exits or someone runs Recover.
  • Old automation shells: automation terminals left from before an orcad restart stay open; only those from the current server are cleaned up.
  • Terminals after a Move: each tab whose shell the Move stopped stays open and restarts as a fresh shell on the managed server, with the same layout (fix(ssh): Move to managed server keeps the host's terminal tabs #26077). If the host stays on the relay, only the shells the Move actually stopped are restarted. A restored tab is titled "Terminal" rather than its old number.
  • Terminal checks: they go through the relays; asking the host's terminal daemon directly is a follow-up.
  • Windows CI timing: some Windows CI checks use fixed ~5 s timing budgets and occasionally flake. Scaling them is a follow-up.
  • Leftover relay shells after a rollback and re-upgrade: shells a rolled-back release opened on an already-managed host keep running, and this build can't list or close them yet (needs a stop path for managed hosts).
  • Rolling back to the release: terminals running on the managed server show as "Reconnecting" in the older release, and the server as offline. This is the expected one-way conversion and needs a release note.
  • The relay is not deleted yet: it stays as the fallback for older desktops, rollback, live relay terminals and unsupported hosts. Deleting it and the one-time migration code is a separate cleanup after rollout, once telemetry shows the fallback is rare.
  • Not in Phase 3:
    • moving live terminals;
    • moving a host back off orcad;
    • an orcad updater for packaged macOS orca serve;
    • an end-to-end automation parity test on orcad, and a real-agent completion check on a real host.

Agent skill upstream boundary

  • Not applicable, or this change follows docs/reference/agent-skill-sharing-upstream-boundary.md and copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.

Notes

Security, cross-platform (Linux, macOS and Windows hosts), SSH, folder workspaces, mobile and backward compatibility are covered above. Do not merge without the project owner's explicit approval.

m4air and others added 20 commits October 2, 2026 00:52
…24525)

* feat(orcad): Windows remote primitives for managed orcad hosts (W1)

* refactor(orcad): run Windows host ops as node.exe with plain argv, no PowerShell hop

* fix(orcad): refuse secret-shaped names on the breakaway launcher's --env

* fix(orcad): name the secret env guard for its role

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
…ged server (#16741 T6-9) (#24521)

* feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9)

The destination half of a catalog migration: an orcad stages a T6-7 manifest
(repositories, project groups, folder workspaces, dormant session, client,
automation and worktree metadata, retired names, scrollback snapshots) with
exclusive claims, then commits it with a receipt so a retried commit returns
the same receipt and never imports twice. Served as orcad.migration.* runtime
RPC behind the orcad.migration-catalog.v1 capability; the client refuses a
host without it or with method-not-found, and any other failure is left for
the caller to recheck. Dormant only: no live PTY projection. Inert on the
desktop until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(rpc): catalog the orcad.migration params in the shared contract; name the catalog-import install target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… on proven terminal exit (#16741 T8-c1+c2) (#24522)

A migration from a relay-hosted SSH target into a managed orcad now starts with
a journal in its own sidecar directory, then the target's managed-owner fence,
then a profile flush, before any remote call. A fence with no journal is
unverifiable and never released; a journal whose fence is gone is stale and
grants nothing; an unreadable journal fails closed. The fence requires every
terminal the target ever leased to be proven exited, checked before the fence
(with the relay's process list) and again under it. Same-owner claims now need
the durable record that explains them. Inert until T8-c4/T6-10.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(orcad): run orcad itself on Windows hosts (W2)

* test(orcad): load the ConPTY smoke's addon from out/orcad so the temp slot can be removed

* test(orcad): skip the foreign-uid lock case when running as root

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
The untransferred-dependency census counted activeConnectionIdsAtShutdown naming the
target as workspace-session state. The renderer rewrites that list on every connection
change, so merely connecting to an empty host blocked the move. It is a reconnect hint;
the remote work it can stand for is counted on its own. The empty-target claim check
likewise ignores global-field copies inside the host's session partition.

Co-authored-by: m4air <m4air@Mac.localdomain>
…stination (#16741 T8-c3) (#24523)

* feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3)

The coordinator re-checks before every stage and commit that the fenced source
still exports the journaled manifest, carries no untransferable state and
started no terminal. A lost answer is re-read from the destination's catalog
state; only a committed read whose receipt matches the journal advances it,
and the journal is on disk before anything returns. Abort releases the fence
only on proof the destination holds nothing, or on an unsupported destination
before anything was staged, and never once the destination committed. Codes
against T6-9's catalog client through an injected interface. Inert until
T8-c4/T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(ssh): import the T6-9 client's unsupported refusal instead of mirroring it

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s (W3) (#24563)

* feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3)

* test(orcad): exhaustive op switch in the Windows lifecycle fake

* test(ssh): narrow the Windows host-cell descriptor by lane before building a relay cell

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
…_RUNTIME=orcad (T6-11) (#24608)

* feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11)

* fix(serve): keep orcad selection app-side and wait out Windows temp cleanup

* refactor(orcad): move the data-root privacy check out of the instance lock

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
…ad serve switch (#24619)

* test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch

* ci(e2e): install ripgrep for the serve mode-switch job's window-manager wait

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
…through a journaled migration (#16741 T8-c4) (#24562)

* feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4)

The conversion entry resumes or takes the fence, deploys and pairs the
managed server into it, marks the server as migrated, then stages and commits
the dormant catalog. Every step is keyed by the journal, so a repeat after a
crash, deferral or lost reply resumes the same migration. Status reports an
unfinished migration, and rollback is refused while one runs or when the
rollback snapshot predates the migrated catalog. Inert until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* style: oxfmt the c4 conversion and maintenance files

* fix(ssh): name the fake migration destination's type so declarations stay portable

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…W4) (#24570)

* feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4)

* fix(orcad): accept a managed stop request whose lock path is spelled with Windows client separators

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
…ven commit (#16741 T8-c5) (#24565)

* feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5)

Once the journal records destination-committed, the source profile drops the
manifest's repositories, folder workspaces and unreferenced project groups,
its dormant session, automation, client and worktree state, and the leases
the fence proved exited. The profile flushes, the retirement is verified,
the journal moves to source-retired and compacts once the server matches.
A retry after any crash repeats idempotent work. The fenced target stays: it
carries the managed server's tunnel. Conversion now ends retired. Inert until
T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): retirement drops the migrated host from the reconnect hint

The census no longer treats activeConnectionIdsAtShutdown as untransferable (#24609), so
retirement must remove the target from it; otherwise a restart dials a host that is now a
managed server.

* style: oxfmt the c5 conversion file

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d (W5 part 1) (#24579)

* feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1)

* test(orcad): start the Windows lane's exec spy after the relay gate prelude restores its own

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
…an experimental setting (#16741 T6-6 + T6-10 UI) (#24590)

* feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI)

Adds a Managed servers section under Remote servers (deploy an empty server,
status with deferred-update and migration states, update, rollback, recover,
stop and cancel-stop, and SSH access for paired servers), and a Move to managed
server action on connected macOS and Linux SSH hosts with a preflight summary,
a terminals-closed confirmation and a resumable progress view. Main wires the
conversion to the relay's process list, the direct session and the T6-9 catalog
client. Everything is hidden until the new experimental setting is turned on,
and Windows SSH hosts are never offered. Merges the T6-9 branch (#24521) until
it lands on the integration branch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): align managed-server form controls and name the section the setting reveals

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): ask the relay with an absolute deadline via the W5 terminal-gate lister

The conversion wiring passed a relative 10 s as listProcesses' deadlineMs, which the
provider reads as an absolute time, so every relay inventory timed out after 1 ms and
the terminal gate could never prove exit. Adopt #24579's lister verbatim so the stacks
merge cleanly.

* feat(settings): name blocking saved state in plain, localized words

The move preview listed internal dependency ids such as workspace-session; each kind
now has its own catalog entry.

* feat(settings): offer managed servers and the move on Windows SSH hosts

W1-W5 are on the integration branch, so a Windows relay-hosted host can deploy, convert
and retire like a POSIX one. The move still waits for a connected relay that reported
its platform.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ssions matrix (#24865)

* test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix

* test(ci): expect the orcad Windows host cells in the SSH Windows hosts workflow

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
… their temp profiles (#24871)

Co-authored-by: m4air <m4air@Mac.localdomain>
…yload, all inactive (#24867)

The nudge request Orca already polls may now carry an optional versioned rollout block
naming the Node runtime flips. A typed reader resolves each flip with kill-switch, version
range and install-id-bucketed percent semantics, falling back to the baked value (every flip
inactive) when the block is absent, invalid or never read. No consumer reads it yet.

Co-authored-by: m4air <m4air@Mac.localdomain>
…iable or failed runtime checks (#24866)

ssh_remote_runtime_resolved dropped the self-test's security_software refusal to 'none' and
sent nothing when a self-test was unverifiable or failed, because those attempts throw before
a rung settles. Add the refusal value, self_test 'unverifiable', and an outcome field
(resolved | unverifiable | failed) deduplicated per host and outcome per session.

Co-authored-by: m4air <m4air@Mac.localdomain>
OrcaWin and others added 3 commits October 2, 2026 14:37
…le (#24884)

global-settings-types.ts sits at the 300-line max-lines ceiling; merging main's two new
agent-state-rules settings with experimentalManagedServers put it at 301. The worktree
visibility defaults type moves next to the other visibility types and is re-exported so its
38 importers are unchanged.

Co-authored-by: m4air <m4air@Mac.localdomain>
# Conflicts:
#	config/ci/windows-ssh-provider/preview-ssh/prove-preview-openssh.ps1
#	src/main/ipc/parcel-watcher-process-supervisor.ts
… canary into a child slot (#24887)

* refactor(watcher): move the supervisor's child, terminating child and canary into a child slot

* fix(watcher,runtime): take the child slot's child type from the shared wrapper, and stub main's title-display clear in the projection test

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
@OrcaWin OrcaWin closed this Oct 3, 2026
@OrcaWin OrcaWin reopened this Oct 3, 2026
…ws builds (#24969)

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
@OrcaWin
OrcaWin deployed to adhoc-mac-build October 3, 2026 07:45 — with GitHub Actions Active
@OrcaWin
OrcaWin deployed to adhoc-mac-build October 3, 2026 07:54 — with GitHub Actions Active
…cad versions after each managed deploy (#24973)

* feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy

* fix(orcad): keep the in-use slot plus the two most recent others, and prove same-version reuse survives eviction

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
m4air and others added 3 commits October 7, 2026 01:25
…3-sync5

# Conflicts:
#	config/scripts/run-node-server-tests.mjs
…#26087)

* fix(ssh): reclaim this desktop's own exited lock on Windows hosts too

The relaunch after a quit mid-update now frees the activation fence and the
version-dir install lock on a Windows SSH host the same way it does on POSIX,
instead of waiting out the 20-minute stale window. The host script ages the
lock only when its token belongs to a desktop process proven exited, it has
been quiet for three heartbeats, and (for the fence) no state mutation is
live, where a mutation holder counts as gone only by pid plus creation time.

* fix(ssh): take an exited holder's lock only through the steal arbitration

Review found the reclaim backdated the lock by path after checking it, so a
live successor that replaced the lock in between could be aged and then
stolen, and an interrupted or failed restore left it aged for good.

The exited-holder check is now read-only. The steal command itself accepts
the proven token and, inside its steal claim and identity recheck, also takes
a lock whose owner file still names that token and that has been quiet for
three heartbeats. Nothing is written to a lock before the steal owns it.
POSIX uses the same path.

* fix(ssh): never take an exited holder's fence while a state mutation can start

Review round 2 found the fence's live-mutation guard ran only in the read-only
proof, so a mutation admitted after the proof, or one whose first heartbeat
landed after the steal sampled the fence's age, kept running under a fence
the steal had replaced.

For the fence, the steal now takes the state-mutation lock inside its claim
(mkdir on POSIX, the exclusive owner.json on Windows) and holds it until the
takeover is done; it refuses when any mutation lock exists. Holding it, it
rereads the owner and only then re-samples the fence identity. A mutation now
rechecks its fence token right after it takes the mutation lock and stops with
the fence-lost marker if it changed. The Windows proof also falls back to the
stale window when its command line would not fit cmd.exe.

* fix(ssh): record the exited-owner steal as a real mutation-lock holder

Review round 3 found the POSIX steal held the state-mutation lock as an empty
directory, which a mutation reclaims after a minute without any liveness
check; a steal stalled that long lost its exclusion and could replace the
fence under a running mutation.

The steal now writes its pid (and group, under the same rule) with the
mutation's own noclobber owner writer, so only proof of its exit frees the
lock, and it removes the lock only while the lock still names it. On
Windows the owner record is moved into place whole, so it never exists
empty, and is removed only while it still names the steal's pid.

* refactor(ssh): keep the relay lock commands off the orcad host-script graph

The mutation-lock owner writers moved into a leaf module, so the relay's
install-lock commands no longer import orcad-state-snapshot and, through it,
the Windows host script, orcad-instance-lock and the daemon process query.
Those modules evaluate imports at load time that existing suites mock
partially. No behavior change.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
@OrcaWin
OrcaWin merged commit 5cafefe into main Oct 7, 2026
146 of 148 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant