Transparent stdio multiplexer that lets multiple Claude Code sessions share a single MCP server process.
One line change in .mcp.json — no other configuration required.
Each Claude Code session spawns its own copy of every configured MCP server (stdio transport). With 4 parallel sessions and 12 servers, that is 48 node/Python processes consuming roughly 4.8 GB of RAM. Most MCP servers are stateless — they don't need per-session isolation.
mcp-mux consists of two components: a thin shim (the binary CC invokes) and a long-lived daemon that owns upstream processes. Shims connect to the daemon via IPC; the daemon spawns and manages upstream servers on behalf of all shims.
graph TB
subgraph "CC Sessions"
CC1[CC Session 1]
CC2[CC Session 2]
CC3[CC Session 3]
end
subgraph "mcp-mux Daemon"
D[mcp-muxd]
O1[Owner: engram]
O2[Owner: tavily]
O3[Owner: aimux]
R[Reaper/GC]
end
subgraph "Upstream Servers"
U1[engram]
U2[tavily]
U3[aimux]
end
CC1 -->|"stdio → shim → IPC"| O1
CC2 -->|"stdio → shim → IPC"| O1
CC2 -->|"stdio → shim → IPC"| O2
CC3 -->|"stdio → shim → IPC"| O3
O1 -->|stdio| U1
O2 -->|stdio| U2
O3 -->|stdio| U3
Each shim connects to the daemon owner for its upstream. If no daemon is running, the shim auto-starts one. If no owner exists for a given server, the daemon spawns it.
Result: one upstream process per server instead of N — approximately 3x memory reduction.
The prepared binary target is v0.31.0. After publication, use the v0.31.0 release. The build commands below are source builds, not delivered-artifact proof.
1. Build
# Linux / macOS
go build -o mcp-mux ./cmd/mcp-mux
# Windows
go build -o mcp-mux.exe ./cmd/mcp-muxPlace the binary somewhere on your PATH, or reference it by absolute path in .mcp.json.
2. Configure
Take any MCP server entry in .mcp.json and move the command into args[0], replacing
command with mcp-mux:
Before:
{
"mcpServers": {
"engram": {
"command": "uvx",
"args": ["engram-mcp-server", "--db", "/data/engram.db"]
}
}
}After:
{
"mcpServers": {
"engram": {
"command": "mcp-mux",
"args": ["uvx", "engram-mcp-server", "--db", "/data/engram.db"]
}
}
}3. Verify
mcp-mux statusOn the next CC session start, mcp-mux intercepts the stdio channel, connects to (or starts) the daemon, and proxies all MCP traffic transparently.
| Mode | Behavior | Use When |
|---|---|---|
shared (default) |
One upstream serves all sessions. Responses to initialize, tools/list, prompts/list, and resources/list are cached and replayed without a round-trip. |
Stateless servers: search, docs, LLM proxy. |
isolated |
Each session gets its own upstream process. | Per-session state: browser automation, SSH, editor buffers. |
session-aware |
One upstream; sessions identified by injected _meta.muxSessionId. |
Stateful servers that can partition in-process state by session key. |
Override mode for a specific server:
# Force isolation for one invocation
MCP_MUX_ISOLATED=1 mcp-mux uvx my-server
# CLI flag (equivalent)
mcp-mux --isolated uvx my-serverUse R1 only when the host and upstream are both known to speak MCP 2026-07-28. It is an additive opt-in. Omitting --mcp-protocol=2026-07-28 preserves the legacy quick start and default behavior. R1 does not translate legacy and modern traffic.
Put --mcp-protocol=2026-07-28 before the upstream command:
mcp-mux --mcp-protocol=2026-07-28 my-modern-server --stdioFor example, configure the server in .mcp.json as follows:
{
"mcpServers": {
"modern-server": {
"command": "mcp-mux",
"args": [
"--mcp-protocol=2026-07-28",
"my-modern-server",
"--stdio"
]
}
}
}R1 forces isolation. It disables response caching, discovery templates, and replay. There is no automatic fallback to the legacy protocol or a shared mode. Server log notifications stay request-scoped and require the request to opt in with _meta.io.modelcontextprotocol/logLevel; mcp-mux does not broadcast them.
After an owner or transport loss, mcp-mux returns errors for in-flight modern requests. The host must issue fresh retries and a fresh subscriptions/listen; R1 does not replay requests or subscriptions.
To inspect the R1 policy, run mcp-mux status and look for these four fields:
protocol_era: "2026-07-28"sharing_policy: "forced-isolated"cache_policy: "off"lifecycle_policy: "r1-quarantine"
R1 is separate from the historical --stateless flag. --stateless changes legacy server identity and sharing. It does not select an MCP protocol era or make R1 shareable.
To roll back an R1 configuration:
- Remove
--mcp-protocol=2026-07-28from the host configuration to stop new R1 admissions. - Let current R1 owners drain, or remove them with
mux_stopby theserver_idshown inmcp-mux status.mux_stopis the existing control-plane tool for any owner. R1 adds no separate stop path.mcp-mux stop --drain-timeout 30sdrains and stops the whole local daemon. - Start a new host connection with the legacy configuration. Never downgrade a live R1 owner or replay old requests or subscriptions.
A refused, absent, or mismatched era confirmation means modern admission failed. Do not retry the same route as legacy.
When no explicit mode is set, mcp-mux classifies each server automatically using this priority order:
x-muxcapability (highest) — server declaresx-mux.sharingin itsinitializeresponse. Authoritative; overrides all heuristics.- Tool-name heuristics — tools with names matching browser, session, editor, navigate, page, tab, process, document, or snapshot patterns trigger isolation.
- Default —
shared.
flowchart TD
A[Server starts] --> B{x-mux capability\nin initialize response?}
B -->|Yes| C[Use declared mode]
B -->|No| D{Tool names match\nisolation patterns?}
D -->|Yes| E[Isolated]
D -->|No| F[Shared]
If your server is stateless but has tool names that match isolation patterns, add
"x-mux": { "sharing": "shared" } to your initialize capabilities to fix the classification.
In shared mode, the owner intercepts and caches the first response for each of these methods:
initializetools/listprompts/listresources/listresources/templates/list
Subsequent sessions receive the cached response immediately without a round-trip to the upstream.
Cache entries are invalidated when the upstream sends the corresponding *_changed notification
(notifications/tools/list_changed, notifications/prompts/list_changed,
notifications/resources/list_changed).
For initialize, the cache is keyed on protocolVersion. A new client using a different protocol
version bypasses the cache and goes to the upstream directly.
The first time mcp-mux sees a command and security context, it starts the
upstream and performs a synthetic initialize plus tools/list to learn the
server's sharing mode and publish a reusable discovery template.
After that owner is removed, a compatible session can start as a cache-only owner:
- cached
initialize,tools/list, and any captured prompt/resource discovery responses are replayed without starting the upstream process; notifications/initializedis suppressed while no upstream exists;- the first uncached request starts exactly one upstream generation and forwards the original request on the same open host transport after initialization;
mcp-mux statusreportsmaterialization_state: "CACHE_ONLY"withupstream_pid: 0until demand arrives, thenREADYwith the live PID.
Template reuse is fail-closed. A missing, stale, or incompatible template takes one bounded cold/eager path instead of replaying cached frames from the wrong context. Compatibility uses an exact SHA-256 identity of the effective security-relevant environment; isolated templates also require the exact canonical working directory. Raw environment values are not stored in the identity or exposed in status. Persistent owners and callers that explicitly require eager startup continue to materialize without waiting for a request.
For slow-starting servers (serena via uvx ~3s, tavily via npx ~5s), a compatible template lets later host startup complete from cache while deferring that cost until the server is actually used.
The daemon is enabled by default. It starts automatically when the first mcp-mux shim connects and no daemon is running.
Lifecycle:
- Shim connects → daemon starts or is reused.
- A non-persistent initialized shim with no requests or queued work parks its daemon IPC session after 10 minutes without host traffic. New host demand reconnects to the exact owner.
- If no demand arrives for another 30 seconds, the stable launcher parks the
engine process too. The next host frame starts the current active engine,
replays the cached
initializehandshake, and forwards that demand once. - When the last session disconnects, a disposable owner is eligible for the existing 30-second safety-gated zero-session cleanup. The general owner idle timeout remains 10 minutes; isolated owners use their shorter lifecycle rule.
- Servers declaring
x-mux.persistent: truedo not suspend or reap while idle; they stay alive until explicitly stopped or until the daemon exits. - Daemon auto-exits after 5 minutes with no owners and no connected sessions.
MCPMUX_SHIM_IDLE_TIMEOUT and MCPMUX_SHIM_DORMANT_GRACE override the
10-minute and 30-second product-shim stages with Go duration strings. Zero or a
negative value disables that stage; invalid values keep the default. These are
separate from owner cleanup and daemon auto-exit settings.
Serena's web dashboard is configured separately from mux lifecycle. To prevent
it opening automatically, pass --open-web-dashboard false to Serena's
start-mcp-server command or set web_dashboard_open_on_launch: false in
serena_config.yml; this leaves the dashboard active. Set
web_dashboard: false only when the dashboard itself must be disabled; see the
Serena dashboard documentation.
Standalone upstream launches through MCP_MUX_NO_DAEMON=1, MCP_MUX_DAEMON, or
the direct-owner --daemon path are explicitly maintenance_unsupported in
v0.31.0. Keep managed daemon admission enabled; direct execution is not a
maintenance bypass.
v0.31.0 adds optional maintenance-aware managed hold, resume, and renew operations. Use an aware binary, daemon, and shim together. Publication and consumer handoff remain pending; older installed tags do not acquire maintenance support.
Read the exact server_id from local mcp-mux status. Flags follow the identifier:
mcp-mux hold <exact-server-id> --ttl 5m --drain-timeout 10s --json
mcp-mux renew <returned-hold-id> --ttl 5m --json
mcp-mux resume <returned-hold-id> --json
Hold fences new demand before draining already-forwarded requests. The default
drain is 10 seconds; --drain-timeout 0s skips grace but still requires complete
managed tree death. Positive grace uses the same T as TTL and never restarts
on retirement retry. Replace the executable yourself only after a successful
JSON result reports state: HELD, trees_retired: true, and a future expires_at.
The result also contains hold_id, server_id, and drain_deadline.
| State | Meaning |
|---|---|
HOLDING |
Admission is fenced; provisional timing precedes clocked drain/retirement. No replacement grant. |
HELD |
Every scoped managed tree is proven dead; the lease is usable until expiry. |
RETIREMENT_BLOCKED |
Tree death is unproven; replacement is unsafe and admission stays fenced. |
RELEASED |
Durable release permits fresh demand. |
TTL defaults to five minutes, must be positive, and cannot exceed one hour. Renewal uses the exact current unexpired hold ID and sets expiry from serialized renewal acceptance, without changing retirement state or reviving a released/expired lease. CLI TTL and drain durations must be whole milliseconds; sub-millisecond values are rejected rather than rounded down to zero-force retirement. Competing or stale identities cannot replace or clear a lease. Explicit resume and TTL recovery open admission only after proven retirement and durable release. Blocked retirement never clears because time elapsed.
A durable HOLDING seed has provisional timing and never grants replacement.
After its first complete writer acknowledgment, sample T once and persist clocked HOLDING once.
TTL/drain use T; that write, retirement, HELD persistence, and response consume the original window.
Any acquisition write failure retains the seed fence; incomplete recovery is RETIREMENT_BLOCKED, with no expiry/resume.
With no positive caller timeout, neutral control allows 180s plus one drain for hold/restart_owner; CLI/MCP use this default.
Explicit positive library budgets are honored; other commands retain 5s. This finite exchange allowance does not guarantee full-pin/storage completion.
A timeout leaves outcome unknown: a durable lease/restart may remain. Inspect status, with no automatic retry/resume or stop fallback.
Aware shims keep the original host pipes open. Requests received while fenced
return JSON-RPC -32005, message upstream held for update, and
data.error_code: maintenance_held with the original numeric or string ID.
Maintenance wins over cached success. Rejected requests and unfinished work
ended by retirement are not replayed; notifications receive no invented ID.
Fresh legacy demand after release can reach a replacement on the same pipes.
Modern demand uses fresh same-era isolated admission with required per-request
metadata, without legacy bootstrap, cache, or request/subscription restoration.
Scope is the selected owner's finite already-admitted context set, not a host-wide executable lock. Different CWDs, protocol eras, credentials/configuration contexts, engine namespaces, and unmanaged processes are outside that set unless already admitted explicitly. An unrelated process can still lock the file.
MCP mux_hold, mux_resume, and mux_renew use the same local daemon operation.
mux_restart uses daemon-owned restart_owner with the original exact context
and era, not an adapter's ambient credentials. Controlled restart, handoff,
shutdown, downgrade, and idle daemon exit refuse terminally while a fence remains;
launcher and library update helpers cannot substitute shutdown or a successor.
After unplanned daemon loss, aware startup reloads held authority before admission.
Incomplete or unreadable authority fails closed.
Durable authority is the mandatory schema-2 pair ledger.json and
transaction.json, not either file alone. Pending, missing, or invalid pairs
fail closed. A persistence error keeps live admission conservative but does not
guarantee storage rollback. After a finalize error, recovery can accept only a
matching COMMITTED certificate proving earlier acknowledged durable publication;
the failed caller response remains an error. Successful release cannot resurrect
the old lease. See the storage phase contract.
Controlled engine installation, launcher swap, layout/bootstrap mutation, and
active-pointer updates use the existing daemon namespace file lock. Hold-ledger
mutations use the same lock, so a hold cannot race past an activation check.
daemon.CheckMaintenanceForActivation reads persisted authority without changing
it. Status and pure startup inspection neither acquire nor write that lock and
never proactively start a daemon. Activation checks also consult live aware
status; an offline or old endpoint is allowed only when persisted authority is
proven clear under the lock. This is namespace coordination, not a host-wide
file lock or a separate updater lease.
Old daemons return maintenance_unsupported, never a stop/kill/restart fallback.
An aware daemon fences old managed shims' starts, but those shims have no promised
immediate-error or non-replay semantics. Arbitrary old binaries, foreign engines,
manual active-pointer replacement, and direct standalone bypass are unsupported.
Before downgrade, use the current aware binary to durably resume every exact
retired lease or observe its safely committed expiry. Incomplete/blocked authority
prevents downgrade and must be retained, even after TTL. Do not delete either
authority member or perform PID cleanup to reopen admission.
Run the live Windows and Unix replacement proof before release. Its private primary-checkout scratch and actual overwrite evidence supplement focused regressions; a fixture build or unit test is not that proof.
Outside maintenance fences, legacy mcp-mux shims automatically reconnect when the daemon restarts. Modern R1 does not replay the legacy handshake below. This means:
mcp-mux upgradeswitches the active versioned engine without dropping connectionsmcp-mux stop --forcetriggers automatic reconnect within seconds- Daemon crashes are recovered transparently
During reconnect, the shim:
- Detects IPC connection loss (daemon shutdown)
- Drains orphaned in-flight requests — sends spec-compliant JSON-RPC error responses so CC sees explicit failures for pending requests instead of silence (silence on a pending request is what CC's stdio transport tears the connection down over)
- Tries the planned reconnect path first, including reconnect-token refresh
- Starts or waits for the replacement daemon when needed
- Replays cached
initializerequest to warm the replacement owner - Sends
notifications/tools/list_changedso the host can refresh discovery - Flushes any still-valid buffered requests and resumes normal proxy
Reconnect is transport continuity, not request replay. A request already sent
to the lost owner receives one explicit JSON-RPC error with its original id and
is never sent to the successor. Only the cached initialize handshake is
replayed to warm the replacement connection; host frames accepted afterward
are forwarded once.
Reconnect timeout: 30 seconds. If reconnect is still unavailable after that window and a reconnect path exists, the shim does not exit. It enters degraded retry: new client requests receive JSON-RPC errors by their original ids, the parent stdio transport stays open, and normal proxying resumes on that same transport after backend recovery. The shim exits only when the MCP host closes stdin/stdout or when no reconnect path was configured.
Note on keepalives: Earlier versions emitted synthetic
notifications/progresswith amux-reconnectprogress token every 5 s as a keep-alive. That violated the MCP spec (progress tokens must reference a client-issued_meta.progressToken), and Claude Code tore down the stdio transport on the first unknown token — destroying the connection the shim was trying to preserve. The keep-alive was removed in muxcore v0.19.6;drainOrphanedInflightis the spec-compliant replacement.
mcp-mux v0.4.0 introduces a session transport layer that replaces the old lastActiveSessionID
heuristic with deterministic, per-session routing.
When CC spawns a shim, the daemon generates a cryptographic token tied to that spawn's working directory. The shim sends this token as the first line on the IPC connection:
CC → shim → [token\n] → Owner (SessionManager) → upstream
The Owner reads the token, looks up the corresponding Session.Cwd, and binds the IPC connection
to that session. From this point the session identity is authoritative — no heuristics required.
Handshake enforcement (v0.9.10+). The Owner rejects IPC connections with an empty or
unregistered token when daemon mode is active. Rejections are logged at owner level with the peer
PID (no token value) and rate-limited to 10 entries per minute per owner with a suppressed-count
summary. Pre-registered tokens are preserved on rejection, so a legitimate client that closes
mid-handshake can reconnect without forcing the daemon to re-issue a new token. Tokens are 128-bit
(16 random bytes from crypto/rand); entropy failure is fatal.
The SessionManager tracks inflight requests per session. When exactly one session has pending
requests outstanding, response routing is deterministic without needing to inspect message content.
This eliminates spurious mis-routing in high-concurrency scenarios.
roots/list requests from the upstream are forwarded to the active CC session (the one with
pending requests), so the server receives the real workspace roots for that session rather than a
static fallback.
mcp-mux is designed for a single-user local trust boundary: any process running as the same OS user is implicitly trusted. Two layered defenses protect against same-machine impersonation on shared Unix hosts:
The Owner acceptLoop rejects IPC connections with an empty or unregistered token (daemon mode).
Combined with 128-bit crypto/rand tokens and single-use Bind semantics, this closes the only
application-layer impersonation gap on the data socket.
All Unix domain sockets created by ipc.Listen and the daemon control socket go through the
muxcore/sockperm package, which applies syscall.Umask(0177) under a package-level mutex — the
socket file lands with mode 0600 and is only accessible to the owner UID. On Windows, AF_UNIX
sockets inherit the creating process's default DACL (owner + LocalSystem), so no umask equivalent
is needed and the package is a documented no-op.
- Malware running under the same user account. A process with your UID can still connect to
your 0600 control socket and issue its own
spawnrequest to obtain a fresh pre-registered token. Treat the control socket as trusted to everything running as you. - Network-level adversaries. mcp-mux uses Unix sockets / Windows AF_UNIX only — there is no TCP listener. Remote attack surface is zero.
- Upstream MCP servers themselves. mcp-mux is a transparent proxy; if an upstream server runs
exec.Commandon attacker-controlled input, mcp-mux doesn't rewrite or sanitize that.
For shared-machine Unix hosts (multiple login users), mcp-mux v0.9.10 and later is safe for the
cross-user boundary — the 0600 permission prevents a different user from connect()-ing, and the
token handshake rejects same-user probe attempts that haven't received a pre-registered token from
the daemon.
# Show all running upstream instances (PID, sessions, classification, cache state)
mcp-mux status
# Stop all running instances and the daemon
mcp-mux stop [--drain-timeout 30s] [--force]
# Versioned engine upgrade (see section below)
mcp-mux upgrade
# Start a detached daemon process (normally auto-started by shims)
mcp-mux daemon
# Run as control-plane MCP server (exposes mux_list / mux_stop / mux_restart tools)
mcp-mux serveUpgrading the mcp-mux binary while sessions are active is safe and fast, but
ordinary upgrades do not force live host transports through a daemon restart.
The configured mcp-mux executable is a stable launcher and stdio anchor; the
runtime code lives behind the versioned engine pointer.
This section describes the mcp-mux product updater. If you embed muxcore
in another product, start from muxcore/README.md and choose that product's
own launcher, engine path, staging name, and status/update surface instead of
copying the mcp-mux.exe~ / mcp-mux.versions layout blindly.
# One-command upgrade with safe restart/defer semantics
go build -o mcp-mux.exe~ ./cmd/mcp-mux && mcp-mux upgrade --restartWhat happens:
- The stable
mcp-muxlauncher is left in place; it is not renamed while live shims hold it. - The pending
mcp-mux.exe~engine is copied or moved intomcp-mux.versions/<hash>/mcp-mux-engine.exe. mcp-mux.versions/active.txtis switched to the new engine path.- If the daemon has live sessions, the daemon restart is deferred. Existing host stdio transports stay attached to their current daemon/engine boundary, and new shims use the new active engine pointer.
- If zero live sessions can be proven,
--restartmay perform graceful restart: the daemon serializes state snapshot (cached init/tools/prompts/resources responses, classification, session metadata), shuts down, and starts the successor daemon from the new engine. - Current-generation shims reconnect to the successor, refresh reconnect tokens, and resume against restored owners with pre-populated caches.
The launcher indirection is intentional on Windows: live mcp-mux.exe shim and daemon
processes keep the executable image locked, so a self-rename of the configured binary is
not a reliable update primitive. Versioned engines avoid that lock; old processes keep
running from their old engine path while new shims use the active engine pointer.
This has one bootstrap boundary: processes that were already launched by an
older, pre-supervisor mcp-mux serve cannot be retrofitted in-place. Move the
host configuration onto the stable launcher once; after that, ordinary engine
updates preserve the host-facing stdio boundary by keeping the launcher stable
and deferring daemon restart when live sessions are present.
Without --restart (active pointer switch only):
mcp-mux upgradeThe daemon keeps running with its current engine. New shim processes use the new active engine. The daemon updates on next natural restart.
When a same-protocol graceful restart is safe and actually runs, it preserves:
- The upstream process tree — a same-v2 handoff retains it after successor adoption; a restart with live sessions is deferred instead
- Cached MCP responses (init, tools, prompts, resources)
- Server classification (shared/isolated/session-aware)
- Session metadata (cwd, env)
- Reconnect-token history, so live shims can refresh without fallback-spawning an owner during planned restart.
Only the daemon restarts — upstreams are reattached via FD passing (Unix SCM_RIGHTS, Windows DuplicateHandle). See the next section for the lifecycle contract.
Handoff protocol v2 keeps an upstream process tree alive across a planned same-v2 daemon restart only after the successor adopts both stdio and the tree authority. Ordinary reconnect after an owner loss preserves the host transport but does not replay already-sent requests.
| Trigger | Pre-v0.21.0 | Current contract |
|---|---|---|
mcp-mux upgrade --restart with live sessions |
Upstream killed + respawned, in-flight requests dropped | Daemon restart deferred; existing stdio transports remain on the current daemon while new shims use the new engine pointer |
mcp-mux upgrade --restart with zero live sessions |
Upstream killed + respawned, in-flight requests dropped | Same-v2 handoff retains the upstream tree; the first v1-to-v2 restart takes one snapshot-backed respawn |
| Daemon/owner loss | Upstream lifecycle was leader-oriented | Recovery is demand-driven; abandoned generations are cleaned as full trees and already-sent requests receive explicit errors instead of replay |
mux_restart <sid> (operator-initiated) |
Hard kill | Daemon-owned restart_owner preserves the exact launch context and era, with 30s drain by default. force: true skips drain only; maintenance-held targets refuse terminally without stop/spawn fallback. |
| Reaper idle-eviction | Hard SIGKILL | Soft-close: 30s stdin drain → SIGTERM only after timeout |
*Unix (Linux, macOS, BSD):
- Upstream spawns with
Setpgid=true— the kernel places the child in its own process group. - Planned restart: old daemon opens a Unix domain socket, successor daemon connects with a 128-bit shared token, FDs (stdin, stdout, stderr) transfer via the SCM_RIGHTS ancillary control message.
- Cleanup targets the process group, including descendants that outlive the leader or inherit its stdio. Planned v2 handoff retains the PGID authority until the successor's final adoption acknowledgment.
Windows:
- Each upstream starts suspended and is placed in its own anonymous Job Object
with
JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSEbefore it can execute. - Planned restart: successor is spawned with a named-pipe address; handles are duplicated via
DuplicateHandlewithDUPLICATE_SAME_ACCESS, including the Job authority. - The predecessor retains its Job lease until the successor commits adoption; abort or final authority loss terminates the whole tree.
Old daemon → successor handshake is JSON-over-socket with a mandatory protocol_version: 2
field on every message:
Hello ──(token, source_pid)──>
<──(protocol_version check, refs list)── Ready
FdTransfer ──(server_id, stdio + tree-authority metadata)──>
<──(SCM_RIGHTS / DuplicateHandle)── AckTransfer (ok/aborted)
...repeat per upstream...
Done ──(transferred, aborted lists)──>
<──(accepted + aborted partition)── HandoffAck
- Token auth (FR-11): constant-time compare, 128-bit random, 0600 file.
- Per-upstream atomicity (FR-7): receipt does not detach the predecessor. Each tree commits only if the final acknowledgment accepts its server id; every other prepared tree aborts and is eligible for snapshot respawn.
- 30s accept + total timeout on both sides.
- Version skew (FR-3): negotiation happens before any owner detaches. A
mismatched
protocol_version, token mismatch, listener timeout, or handoff successor start failure aborts the restart, releases restart pins, and leaves the predecessor serving.
Snapshot fallback is authorized only after exact version/token Hello and owner detach. On a later receipt or final-ack failure, the daemon:
- proves the failed handoff successor has exited;
- aborts the prepared transfer, permanently closing the detached predecessor owner and terminating its process tree;
- rewrites the pinned snapshot from the retained logical generation metadata;
- pre-starts exactly one clean snapshot successor; and
- only then returns the post-response predecessor shutdown callback.
No-process/cache-only restarts likewise pre-start one snapshot successor before
predecessor shutdown. A pre-detach failure logs handoff.abort and does not
spawn fallback. A successful post-detach recovery logs handoff.fallback;
failure to prove successor exit, rewrite the snapshot, or start the backup logs
handoff.fallback_blocked and remains fail-closed.
The snapshot successor eagerly restores upstreams; drainOrphanedInflight
returns JSON-RPC errors by original id to in-flight callers. It does not replay
those requests.
New counters in mux_list / HandleStatus:
| Counter | Meaning |
|---|---|
handoff_attempted |
Total HandleGracefulRestart invocations that entered the handoff path |
handoff_transferred |
Successfully handed-off upstreams across all handoffs |
handoff_aborted |
Upstreams that fell back per-upstream (FR-7) while siblings succeeded |
handoff_fallback |
Whole-handoff failures that took the FR-8 respawn path |
Structured log markers: handoff.start, handoff.upstream.transferred, handoff.complete,
handoff.fallback, handoff.receive.{start,complete,fail}.
No consumer code change is required. The first restart from a v1 handoff binary to v0.27.0 rejects live transfer before detach and uses one bounded snapshot-backed respawn. Subsequent v2-to-v2 planned restarts retain stdio and tree authority transactionally. Rollback across the same v2/v1 boundary uses the bounded respawn again; do not force mixed-version live handoff.
Snapshot back-compat: v0.20.x OwnerSnapshot files load without errors; new fields
(UpstreamPID, HandoffSocketPath, SpawnPgid) are omitempty and default to zero on
old snapshots.
- Protocol transition boundary: the first handoff v1 → v2 restart, and a rollback across that same boundary, reject live transfer before owner detach and take one bounded snapshot-backed shutdown-and-respawn path. Same-v2 restarts retain stdio and full-tree authority only after final successor adoption.
- Per-upstream 30s transfer bound: upstreams that don't drain within 30s fall back to respawn for that entry only.
- macOS launchd cross-parentage: verified via CI; spawns outside the mcp-mux process tree inherit correctly.
- Windows process tree: each upstream is governed by a Job Object with
JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE; final authority loss intentionally terminates descendants, including those that outlive their leader.
# Unix
scripts/verify-handoff.sh
# Windows
scripts\verify-handoff.ps1The script spawns a test daemon, triggers upgrade --restart, asserts all upstream PIDs
survive across the restart, and reports any dropped FDs.
Developers embedding muxcore in another Go MCP server should start with
muxcore/README.md. The recommended integration path is engine.New +
engine.Run: muxcore then owns daemon/client/proxy mode selection, token
handshake, reconnect, session routing, snapshot restore, and graceful restart.
The lower-level handoff functions in muxcore/daemon remain public for custom
daemon supervisors, but ordinary consumers should not build their own shim or
connect directly to owner IPC paths. Doing so bypasses the daemon-minted session
token path and will be rejected by daemon-managed owners.
See muxcore/README.md for the consumer checklist, required engine.Config
fields, upgrade/launcher contract, guardrails, and low-level handoff API. The
guide also states the muxcore API design rule: infer safe behavior in engine
when possible, and fail early with an actionable error when the consumer wiring
is ambiguous.
- Spec:
.agent/specs/upstream-survives-daemon-restart/spec.md - Engram:
#109(arc resolution),#130(public API export for aimux-class consumers)
All configuration is via environment variables. No config file is required.
| Variable | Default | Description |
|---|---|---|
MCP_MUX_NO_DAEMON |
0 |
Legacy setting; 1 is unsupported and fails closed with maintenance_unsupported. Keep daemon-managed admission enabled; this is not a daemon bypass. |
MCP_MUX_ISOLATED |
0 |
Set to 1 to force isolated mode for this invocation |
MCP_MUX_STATELESS |
0 |
Set to 1 to ignore cwd in server identity hash (enables global deduplication) |
MCPMUX_SHIM_IDLE_TIMEOUT |
10m |
Safe host-idle period before a non-persistent shim parks its daemon IPC session; zero or negative disables |
MCPMUX_SHIM_DORMANT_GRACE |
30s |
Exact-owner reconnect window before a supervised launcher becomes dormant; zero or negative disables |
MCPMUX_LAUNCHER_DORMANT_LEASE |
disabled | Explicit opt-in to exit a dormant launcher after this additional no-demand lease. Use only with a host proven to relaunch after transport closure. |
MCP_MUX_OWNER_IDLE |
10m |
General owner idle timeout; overridden per owner by x-mux.idleTimeout |
MCP_MUX_GRACE |
10m |
Legacy alias used only when MCP_MUX_OWNER_IDLE is unset |
MCP_MUX_IDLE_TIMEOUT |
5m |
Daemon auto-exit after this period with no activity |
MCP stdio has no standard logical-completion signal: the host owns termination by
closing child stdin, while a server may send notifications or requests during an
otherwise silent interval. mcp-mux therefore exits on host EOF but does not
kill a silent live stdio transport by default. A capable launcher parks its
engine after the shim idle/grace sequence and retains a small launcher stub for
later demand. MCPMUX_LAUNCHER_DORMANT_LEASE bounds the complete disposable
launcher/engine tree only as an explicit host-compatibility opt-in.
The installed stable launcher and active versioned engine are distinct binaries and may differ byte-for-byte. Private dormant frames require protocol-v2 target-bound bilateral attestation: before spawning the engine, the launcher binds a one-shot current-user local IPC endpoint, then binds the exact PID returned by the child start. The launcher accepts the fixed proof only from that client PID; the engine independently requires the endpoint server PID to be its OS direct parent and completes the fixed request/response exchange on that side channel, never on host stdio. The provider-derived version-store layout, active-engine pointer, and direct-parent executable path must also match the installed stable launcher. Custom or copied engine paths fail closed.
Forwarding the endpoint environment through a v0.27.0-or-older launcher does not transfer capability: the endpoint server remains an ancestor rather than the engine's direct parent. That running session stays fail-closed and receives no private dormant frames. The verified child may bootstrap the stable launcher for future invocations with the rollback-capable two-rename swap; restart only a host/session still running under the pre-v2 launcher before expecting launcher dormancy or a lease.
Customer-mode proof remains required for host relaunch behavior and live Windows
executable swapping. Windows verifies the named-pipe server with
GetNamedPipeServerProcessId, Linux uses SO_PEERCRED, and macOS uses
LOCAL_PEERPID; the existing parent-image checks use the Windows process image,
/proc/<ppid>/exe, and kern.procargs2 respectively. Unsupported platforms,
including BSD targets without both proofs, fail closed and never emit private
dormant frames.
The prepared current library target is muxcore/v0.31.0, which adds optional
managed maintenance to the existing stable stdio supervisor and explicit native
MCP 2026-07-28 route. After publication and Go proxy tag resolution, pin:
go get github.com/thebtf/mcp-mux/muxcore@v0.31.0Ordinary legacy engine.New consumers require no source changes. Use
control.SendMaintenance only when adopting maintenance. Select
engine.Config.ProtocolPolicy explicitly for a known same-era modern route;
legacy remains the zero-value default. Supervisor users should keep using
supervisor.Run, supervisor.StartCommand, supervisor.StartWithFallback,
supervisor.ProtocolV2, and the attestation package rather than copying the
product adapter or private wire codec. Roll back to muxcore/v0.30.0 or a
compatible prior binary only after exact retired leases are durably resumed or
safely expired through the current aware version. Preserve incomplete/blocked
authority and retire modern owners through quarantine, never live conversion.
The shared daemon is owned by the stable launcher rather than by any supervised engine generation. The launcher prepares it before starting a child; on Windows the daemon therefore remains outside the child's KillOnJobClose Job Object. A supervised child never spawns that daemon inside its own process tree and exits back to the stable launcher if the daemon must be recreated, preserving the host-facing stdio pipe.
mcp-mux serve exposes an MCP server on stdio with management tools. Add it to .mcp.json like
any other server:
{
"mcpServers": {
"mcp-mux": {
"command": "mcp-mux",
"args": ["serve"]
}
}
}Tools:
| Tool | Description |
|---|---|
mux_engines |
Lists opted-in native muxcore daemon engines registered on this host. Each descriptor is advisory and is verified by daemon status before being marked healthy. Stale or mismatched descriptors are labeled instead of mixed into owner lists. duplicate means more than one healthy descriptor advertises the same engine name; stale leftovers do not make a healthy daemon duplicate. |
mux_prune_engines |
Dry-run by default. Lists or removes stale / invalid native muxcore registry descriptor files after the same verification used by mux_engines. This is registry garbage collection only: it never stops processes, owners, daemon control sockets, or live native muxcore products. |
mux_list |
Returns running instances for the current project inside this mcp-mux daemon namespace (filtered by caller's cwd). Pass all: true to list this daemon's instances across all projects. Pass exact engine_name from mux_engines to query one verified native muxcore engine explicitly. Includes server ID, engine name, PID, downstream session count, pending requests, classification, and cache status. With verbose: true, includes classification source/reason and inflight request details when present. |
mux_stop |
Gracefully drains and stops an instance by server_id. Use force: true for immediate kill. CR-001 scope is current mcp-mux daemon namespace only; it does not stop native registered engines. |
mux_restart |
Uses daemon-owned restart_owner with the original exact launch context and era. Drain defaults to 30s; force: true skips drain only and never bypasses maintenance-held terminal refusal or enables stop/spawn fallback. Without an explicit target, resolves to the caller's session instance (e.g. mux_restart(name: "aimux") targets this project's mux-managed aimux, not a native engine). CR-001 scope is the current namespace only; cross-engine restart is a future opt-in feature. |
Session-scoped control plane:
The control plane is session-aware. Each tool call is resolved in the context of the calling session's working directory:
mux_list— shows only servers owned by the current project by default. Usemux_list(all: true)for a full view across all projects in thismcp-muxdaemon.mux_engines— shows native muxcore products only when they explicitly opt into daemon registry advertisement.mux_prune_engines— shows prune candidates withdry_run: trueby default. Passdry_run: falseonly after reviewing candidates; it removes stale / invalid registry descriptors, not daemon processes.mux_list(engine_name: "aimux")— queries exactly one registered engine after verifying that the descriptor's control socket returns matchingengine_namefrom daemonstatus.mux_restart(name: "aimux")— resolves to the aimux instance started from this project's directory, not a same-named server from a different project.
This prevents accidental cross-project interference when multiple projects use the same server name.
The sessions count in mux_list is a downstream MCP client/shim count, not a
count of visible terminal windows or top-level agent sessions. Some clients run
hidden stdio app-server processes; each one may attach to mcp-mux as a separate
downstream session.
Native muxcore products such as aimux or engram run under their own engine
namespaces when they embed muxcore directly. They do not appear in default
mux_list unless they were launched through the mcp-mux product daemon. If a
native product opts into muxcore daemon registry advertisement, mux_engines
can discover it and mux_list(engine_name: "...") can list that one engine's
owners. Product-native health, sessions, upgrade, and restart surfaces remain
authoritative unless that product later opts into explicit cross-engine
management capabilities.
Prompts:
| Prompt | Description |
|---|---|
mux-guide |
Full reference on architecture, classification, caching, and troubleshooting. |
mux-status-summary |
Calls mux_list and returns a human-readable summary. |
Declare your server's sharing preference in the initialize response capabilities:
{
"protocolVersion": "2025-11-25",
"capabilities": {
"tools": {},
"x-mux": {
"sharing": "shared"
}
}
}For stateless servers that don't depend on the client's working directory, add "stateless": true
to enable global deduplication — one upstream instance regardless of which directory CC is opened
from:
{ "x-mux": { "sharing": "shared", "stateless": true } }For session-aware servers, mcp-mux injects into every request:
_meta.muxSessionId— unique session identifier (format:sess_+ 8 hex chars)_meta.muxCwd— the CC session's project directory (for--project-from-cwdservers)_meta.muxEnv— per-session environment variable diff (API keys, config paths)
{ "x-mux": { "sharing": "session-aware" } }For servers that must stay alive across all session disconnects (e.g., expensive initialization, background indexing), declare persistence:
{ "x-mux": { "sharing": "shared", "persistent": true } }Full protocol specification including implementation examples (TypeScript, Python, Go) and
migration path: docs/mux-protocol.md.
mcp-mux includes a smoke test that validates mux-specific behavior with real upstream servers:
# Basic: verify serena works through mux
SMOKE_CWD=D:/Dev/my-project SMOKE_EXPECT=isolated \
go run testdata/smoke_isolated.go uvx --from git+https://github.com/oraios/serena \
serena start-mcp-server --project-from-cwd
# Isolation check: two projects get separate owners
SMOKE_CWD=D:/Dev/project-a SMOKE_CWD2=D:/Dev/project-b SMOKE_EXPECT=isolated \
go run testdata/smoke_isolated.go uvx --from serena ...
# With tool call
SMOKE_CWD=D:/Dev/my-project SMOKE_TOOL=activate_project \
go run testdata/smoke_isolated.go uvx --from serena ...What it validates (mux behavior, not upstream correctness):
- Spawn via daemon with proactive init
- Classification matches expected mode
- Session isolation: different cwds → different owners for isolated servers
- Init response forwarded correctly through mux
- Optional: tool call forwarded and response returned
# Run tests
go test ./...
# Run vet
go vet ./...
# Build
go build ./cmd/mcp-muxPull requests are welcome. Please ensure go test ./... and go vet ./... pass before submitting.
For significant changes, open an issue first to discuss the approach.
MIT