Repository navigation
Phase 3: every SSH host runs a managed Orca server (orcad), replacing the relay, and orca serve runs on it - #24863
Merged
Merged
Conversation
…24525) * feat(orcad): Windows remote primitives for managed orcad hosts (W1) * refactor(orcad): run Windows host ops as node.exe with plain argv, no PowerShell hop * fix(orcad): refuse secret-shaped names on the breakaway launcher's --env * fix(orcad): name the secret env guard for its role --------- Co-authored-by: m4air <m4air@Mac.localdomain>
…ged server (#16741 T6-9) (#24521) * feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) The destination half of a catalog migration: an orcad stages a T6-7 manifest (repositories, project groups, folder workspaces, dormant session, client, automation and worktree metadata, retired names, scrollback snapshots) with exclusive claims, then commits it with a receipt so a retried commit returns the same receipt and never imports twice. Served as orcad.migration.* runtime RPC behind the orcad.migration-catalog.v1 capability; the client refuses a host without it or with method-not-found, and any other failure is left for the caller to recheck. Dormant only: no live PTY projection. Inert on the desktop until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(rpc): catalog the orcad.migration params in the shared contract; name the catalog-import install target Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… on proven terminal exit (#16741 T8-c1+c2) (#24522) A migration from a relay-hosted SSH target into a managed orcad now starts with a journal in its own sidecar directory, then the target's managed-owner fence, then a profile flush, before any remote call. A fence with no journal is unverifiable and never released; a journal whose fence is gone is stale and grants nothing; an unreadable journal fails closed. The fence requires every terminal the target ever leased to be proven exited, checked before the fence (with the relay's process list) and again under it. Same-owner claims now need the durable record that explains them. Inert until T8-c4/T6-10. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(orcad): run orcad itself on Windows hosts (W2) * test(orcad): load the ConPTY smoke's addon from out/orcad so the temp slot can be removed * test(orcad): skip the foreign-uid lock case when running as root --------- Co-authored-by: m4air <m4air@Mac.localdomain>
The untransferred-dependency census counted activeConnectionIdsAtShutdown naming the target as workspace-session state. The renderer rewrites that list on every connection change, so merely connecting to an empty host blocked the move. It is a reconnect hint; the remote work it can stand for is counted on its own. The empty-target claim check likewise ignores global-field copies inside the host's session partition. Co-authored-by: m4air <m4air@Mac.localdomain>
…stination (#16741 T8-c3) (#24523) * feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) The coordinator re-checks before every stage and commit that the fenced source still exports the journaled manifest, carries no untransferable state and started no terminal. A lost answer is re-read from the destination's catalog state; only a committed read whose receipt matches the journal advances it, and the journal is on disk before anything returns. Abort releases the fence only on proof the destination holds nothing, or on an unsupported destination before anything was staged, and never once the destination committed. Codes against T6-9's catalog client through an injected interface. Inert until T8-c4/T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(ssh): import the T6-9 client's unsupported refusal instead of mirroring it --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s (W3) (#24563) * feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) * test(orcad): exhaustive op switch in the Windows lifecycle fake * test(ssh): narrow the Windows host-cell descriptor by lane before building a relay cell --------- Co-authored-by: m4air <m4air@Mac.localdomain>
…_RUNTIME=orcad (T6-11) (#24608) * feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) * fix(serve): keep orcad selection app-side and wait out Windows temp cleanup * refactor(orcad): move the data-root privacy check out of the instance lock --------- Co-authored-by: m4air <m4air@Mac.localdomain>
…ad serve switch (#24619) * test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch * ci(e2e): install ripgrep for the serve mode-switch job's window-manager wait --------- Co-authored-by: m4air <m4air@Mac.localdomain>
…through a journaled migration (#16741 T8-c4) (#24562) * feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) The conversion entry resumes or takes the fence, deploys and pairs the managed server into it, marks the server as migrated, then stages and commits the dormant catalog. Every step is keyed by the journal, so a repeat after a crash, deferral or lost reply resumes the same migration. Status reports an unfinished migration, and rollback is refused while one runs or when the rollback snapshot predates the migrated catalog. Inert until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * style: oxfmt the c4 conversion and maintenance files * fix(ssh): name the fake migration destination's type so declarations stay portable --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…W4) (#24570) * feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) * fix(orcad): accept a managed stop request whose lock path is spelled with Windows client separators --------- Co-authored-by: m4air <m4air@Mac.localdomain>
…ven commit (#16741 T8-c5) (#24565) * feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) Once the journal records destination-committed, the source profile drops the manifest's repositories, folder workspaces and unreferenced project groups, its dormant session, automation, client and worktree state, and the leases the fence proved exited. The profile flushes, the retirement is verified, the journal moves to source-retired and compacts once the server matches. A retry after any crash repeats idempotent work. The fenced target stays: it carries the managed server's tunnel. Conversion now ends retired. Inert until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(orcad): retirement drops the migrated host from the reconnect hint The census no longer treats activeConnectionIdsAtShutdown as untransferable (#24609), so retirement must remove the target from it; otherwise a restart dials a host that is now a managed server. * style: oxfmt the c5 conversion file --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d (W5 part 1) (#24579) * feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) * test(orcad): start the Windows lane's exec spy after the relay gate prelude restores its own --------- Co-authored-by: m4air <m4air@Mac.localdomain>
…an experimental setting (#16741 T6-6 + T6-10 UI) (#24590) * feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) Adds a Managed servers section under Remote servers (deploy an empty server, status with deferred-update and migration states, update, rollback, recover, stop and cancel-stop, and SSH access for paired servers), and a Move to managed server action on connected macOS and Linux SSH hosts with a preflight summary, a terminals-closed confirmation and a resumable progress view. Main wires the conversion to the relay's process list, the direct session and the T6-9 catalog client. Everything is hidden until the new experimental setting is turned on, and Windows SSH hosts are never offered. Merges the T6-9 branch (#24521) until it lands on the integration branch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(settings): align managed-server form controls and name the section the setting reveals Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(orcad): ask the relay with an absolute deadline via the W5 terminal-gate lister The conversion wiring passed a relative 10 s as listProcesses' deadlineMs, which the provider reads as an absolute time, so every relay inventory timed out after 1 ms and the terminal gate could never prove exit. Adopt #24579's lister verbatim so the stacks merge cleanly. * feat(settings): name blocking saved state in plain, localized words The move preview listed internal dependency ids such as workspace-session; each kind now has its own catalog entry. * feat(settings): offer managed servers and the move on Windows SSH hosts W1-W5 are on the integration branch, so a Windows relay-hosted host can deploy, convert and retire like a POSIX one. The move still waits for a connected relay that reported its platform. --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ssions matrix (#24865) * test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix * test(ci): expect the orcad Windows host cells in the SSH Windows hosts workflow --------- Co-authored-by: m4air <m4air@Mac.localdomain>
… their temp profiles (#24871) Co-authored-by: m4air <m4air@Mac.localdomain>
…yload, all inactive (#24867) The nudge request Orca already polls may now carry an optional versioned rollout block naming the Node runtime flips. A typed reader resolves each flip with kill-switch, version range and install-id-bucketed percent semantics, falling back to the baked value (every flip inactive) when the block is absent, invalid or never read. No consumer reads it yet. Co-authored-by: m4air <m4air@Mac.localdomain>
…iable or failed runtime checks (#24866) ssh_remote_runtime_resolved dropped the self-test's security_software refusal to 'none' and sent nothing when a self-test was unverifiable or failed, because those attempts throw before a rung settles. Add the refusal value, self_test 'unverifiable', and an outcome field (resolved | unverifiable | failed) deduplicated per host and outcome per session. Co-authored-by: m4air <m4air@Mac.localdomain>
Merged
6 of 8 tasks
…le (#24884) global-settings-types.ts sits at the 300-line max-lines ceiling; merging main's two new agent-state-rules settings with experimentalManagedServers put it at 301. The worktree visibility defaults type moves next to the other visibility types and is re-exported so its 38 importers are unchanged. Co-authored-by: m4air <m4air@Mac.localdomain>
# Conflicts: # config/ci/windows-ssh-provider/preview-ssh/prove-preview-openssh.ps1 # src/main/ipc/parcel-watcher-process-supervisor.ts
… canary into a child slot (#24887) * refactor(watcher): move the supervisor's child, terminating child and canary into a child slot * fix(watcher,runtime): take the child slot's child type from the shared wrapper, and stub main's title-display clear in the projection test --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…ws builds (#24969) Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…cad versions after each managed deploy (#24973) * feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy * fix(orcad): keep the in-use slot plus the two most recent others, and prove same-version reuse survives eviction --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
…3-sync5 # Conflicts: # config/scripts/run-node-server-tests.mjs
…#26087) * fix(ssh): reclaim this desktop's own exited lock on Windows hosts too The relaunch after a quit mid-update now frees the activation fence and the version-dir install lock on a Windows SSH host the same way it does on POSIX, instead of waiting out the 20-minute stale window. The host script ages the lock only when its token belongs to a desktop process proven exited, it has been quiet for three heartbeats, and (for the fence) no state mutation is live, where a mutation holder counts as gone only by pid plus creation time. * fix(ssh): take an exited holder's lock only through the steal arbitration Review found the reclaim backdated the lock by path after checking it, so a live successor that replaced the lock in between could be aged and then stolen, and an interrupted or failed restore left it aged for good. The exited-holder check is now read-only. The steal command itself accepts the proven token and, inside its steal claim and identity recheck, also takes a lock whose owner file still names that token and that has been quiet for three heartbeats. Nothing is written to a lock before the steal owns it. POSIX uses the same path. * fix(ssh): never take an exited holder's fence while a state mutation can start Review round 2 found the fence's live-mutation guard ran only in the read-only proof, so a mutation admitted after the proof, or one whose first heartbeat landed after the steal sampled the fence's age, kept running under a fence the steal had replaced. For the fence, the steal now takes the state-mutation lock inside its claim (mkdir on POSIX, the exclusive owner.json on Windows) and holds it until the takeover is done; it refuses when any mutation lock exists. Holding it, it rereads the owner and only then re-samples the fence identity. A mutation now rechecks its fence token right after it takes the mutation lock and stops with the fence-lost marker if it changed. The Windows proof also falls back to the stale window when its command line would not fit cmd.exe. * fix(ssh): record the exited-owner steal as a real mutation-lock holder Review round 3 found the POSIX steal held the state-mutation lock as an empty directory, which a mutation reclaims after a minute without any liveness check; a steal stalled that long lost its exclusion and could replace the fence under a running mutation. The steal now writes its pid (and group, under the same rule) with the mutation's own noclobber owner writer, so only proof of its exit frees the lock, and it removes the lock only while the lock still names it. On Windows the owner record is moved into place whole, so it never exists empty, and is removed only while it still names the steal's pid. * refactor(ssh): keep the relay lock commands off the orcad host-script graph The mutation-lock owner writers moved into a leaf module, so the relay's install-lock commands no longer import orcad-state-snapshot and, through it, the Windows host script, orcad-instance-lock and the daemon process query. Those modules evaluate imports at load time that existing suites mock partially. No behavior change. --------- Co-authored-by: m4air <m4air@Mac.localdomain>
…aWin/node-rt-phase3-sync5
6 of 9 tasks
8 of 9 tasks
8 of 9 tasks
This was referenced Oct 7, 2026
9 tasks done
8 of 9 tasks
This was referenced Oct 9, 2026
8 of 9 tasks
This was referenced Oct 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ELI5
Today Orca installs a small helper (the relay) on every SSH host, and that install often breaks on hosts with the wrong Node version or no build tools. After this PR, every SSH host runs Orca's own server program (orcad) instead. The app ships orcad with a pinned Node and sets it up when you connect.
Hosts you already use move over on their next connect. Orca asks first if that would restart open terminals. If a host can't run orcad, Orca keeps using the relay and says why. There is no setting to turn any of this on.
orca serveruns on orcad too.This was built and reviewed as about 135 smaller PRs on an integration branch. It lands as one change so
mainnever carries half of it.What Changed
Connecting to an SSH host
Before: every connect installed or reused the relay.
After: before any relay starts, the connect picks the host's server.
Reasons a host can't run orcad: unsupported platform, no orcad build in this app, failed native or C-library check, failed runtime self-test, or Windows security software. Orca remembers these per app version. Every other failure retries on the next connect, and a failed attempt always releases its lock.
How the move works: it's journaled and crash-safe: lock the host, set up the server, copy, commit once. An interrupted move finishes or backs out cleanly on the next connect. Each host's saved state stays its own, even when two hosts share a repo id.
Counting terminals honestly: a running shell can't be carried into another server process, so the move needs proof that every relay terminal has exited.
How the app reaches orcad: through an SSH port forward to the port orcad really bound, after checking that whatever answers is this host's orcad, by its encryption key and our login token. orcad gets a new internal id on every restart, so a desktop reconnects fine after another desktop restarted or updated the server. Where the SSH server forbids port forwarding, the app uses a small helper on the normal SSH connection instead (a "stdio bridge").
Old Linux (CentOS 7, glibc 2.17): fully managed. It gets a runtime and file watcher built for old glibc, and its relay fallback uses Orca's own Node, not the host's.
Keeping a managed server running
orca serve. A managed host closes finished automation terminals that nobody opened, keeping the newest 3 per automation. Before an update counts running terminals, it closes those finished, unused automation shells, so they don't block the update.Terminals from the previous Orca version
Before: after an app update, a terminal the old relay still ran was replaced by an empty shell. The real shell kept running out of sight.
After:
Settings and CLI
orca environment status | update | rollback | recover | stop | cancel-stopcommands run the same actions from scripts.recover --accept-changed-state --yesrestores the pre-update snapshot over changed data.Going back to an older Orca
orca serveORCA_SERVE_RUNTIME=electronopts out, and every fallback prints its reason.orca servestarted over SSH.Security
The file orcad writes on a host to say it's ready holds the login token clients connect with. Under a common login umask, other users on a shared host could read it. Now it, the PID file and orcad's log are owner-only, and so are the folders that hold them. A sweep of other files Orca writes on hosts found no other readable secrets.
Disk use
Remote compatibility
Clients and hosts update on their own schedules, so mixed versions are normal.
Telemetry
Each connect sends one anonymous event: which server the host got, how the app reached it, why, and the host's OS, CPU and C library. Conversions, setup failures and the move prompt each get one event too. Every field comes from a fixed list, so no hostnames, paths or log text are sent.
Why
mainwould convert hosts without the downgrade safety, the fallbacks or the recovery.Linked Issue
Phase 3 of the Node runtime migration. Ports #16741, except Bun, live terminal handover, and moving hosts back off orcad.
Visual Proof
The new UI is one status line per host in Settings → SSH Hosts. Examples:
The UI also includes the move toast and dialog. Screenshots of the Managed servers section are in #24590.
Testing
Every constituent PR ran full CI, and ad hoc builds pin every job to one commit. The real-host testing drove the app and the
orcaCLI from ad hoc builds of this branch.Linux and macOS real hosts
Thirteen rounds of testing ran on these hosts:
ok, shell gone)worktree psshows "unverifiable" meanwhile.--forceworks, and rollback with terminals open is refused.recover --accept-changed-state --yesrestores it, and calls work.orca environment stop --yesOf the first 16 real-host bugs (BUG-1 to BUG-16), 15 are verified fixed in a rerun. BUG-12, a dropped network leaving a host stuck, didn't reproduce in 3 tries after its fix. The verified fixes include:
Later reruns found more:
The final real-host run, on the last ad hoc build before landing, passed:
close --allOther fixes that landed late:
Release upgrade and rollback (latest release v1.4.222 ↔ this branch)
Windows real hosts
These ran as CI host cells: inbox and preview OpenSSH, on x64 and arm64, driven with the
orcaCLI.okin 438–654 ms, then the host converts)orca environment rmrefused;orca environment stop --yesdecommissions; reconnect redeploysAll four Windows bugs found in this testing are fixed:
Other evidence
orca servebetween the desktop app and orcad: on Linux, macOS and Windows CI, a terminal survives both directions on one profile, and the right side holds the profile lock each time. On an installed Windows app, the relocated terminal daemon survives desktop → orcad → desktop with the same process.Review loop
Every round reviewed the whole branch against
main. Each finding was checked independently before it counted, and every confirmed finding was fixed in its own PR.The round 9 and pass 6 P1s were a snapshot restore racing a second restore, the readiness file's permissions, and the two-desktop restart identity. The readiness permissions and two-desktop identity are fixed (#25809, #25800). The restore race is fixed in #25811.
Where the loop ended: in round 10, both reviewers found the branch clean. Every fix made after that is clean on its final check.
The review fixes and cleanups removed several thousand lines. The largest single cut is the removal of automatic source retirement (net 1,795 lines).
Last fixes (rounds on each, Opus + Astra, P0/P1 only, all clean at their final heads): #26072 (SSH card status line), #26076 (one host row per machine; 4 P1s fixed, then clean), #26077 (Move keeps tabs; 3 P1s fixed, then clean), #26087 (Windows lock reclaim plus an instance-tied steal on every platform; 4 rounds, then clean).
Main sync:
mainwas last merged into the branch at e547664 (#26108, #26147).Known accepted limits
sshhosts that forbid port forwarding: a host reached only through the systemsshbinary doesn't get the stdio bridge, so it stays on the relay.node <repo>\…is matched; a bare listener shows as external. Local Windows works the same way today.orca environment update --force(ororca terminal close --all). The older server's cleanup rule doesn't release those shells. Hosts first set up by this build aren't affected.orca serve;Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
Security, cross-platform (Linux, macOS and Windows hosts), SSH, folder workspaces, mobile and backward compatibility are covered above. Do not merge without the project owner's explicit approval.