Skip to content
thebtfPublic

About

Transparent stdio multiplexer for MCP servers — share one upstream across multiple Claude Code sessions

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

567 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

English | Русский

CI Go License Platform

mcp-mux

Transparent stdio multiplexer that lets multiple Claude Code sessions share a single MCP server process.

One line change in .mcp.json — no other configuration required.

The Problem

Each Claude Code session spawns its own copy of every configured MCP server (stdio transport). With 4 parallel sessions and 12 servers, that is 48 node/Python processes consuming roughly 4.8 GB of RAM. Most MCP servers are stateless — they don't need per-session isolation.

Architecture

mcp-mux consists of two components: a thin shim (the binary CC invokes) and a long-lived daemon that owns upstream processes. Shims connect to the daemon via IPC; the daemon spawns and manages upstream servers on behalf of all shims.

graph TB
  subgraph "CC Sessions"
    CC1[CC Session 1]
    CC2[CC Session 2]
    CC3[CC Session 3]
  end
  subgraph "mcp-mux Daemon"
    D[mcp-muxd]
    O1[Owner: engram]
    O2[Owner: tavily]
    O3[Owner: aimux]
    R[Reaper/GC]
  end
  subgraph "Upstream Servers"
    U1[engram]
    U2[tavily]
    U3[aimux]
  end
  CC1 -->|"stdio → shim → IPC"| O1
  CC2 -->|"stdio → shim → IPC"| O1
  CC2 -->|"stdio → shim → IPC"| O2
  CC3 -->|"stdio → shim → IPC"| O3
  O1 -->|stdio| U1
  O2 -->|stdio| U2
  O3 -->|stdio| U3
Loading

Each shim connects to the daemon owner for its upstream. If no daemon is running, the shim auto-starts one. If no owner exists for a given server, the daemon spawns it.

Result: one upstream process per server instead of N — approximately 3x memory reduction.

Quick Start

The prepared binary target is v0.31.0. After publication, use the v0.31.0 release. The build commands below are source builds, not delivered-artifact proof.

1. Build

# Linux / macOS
go build -o mcp-mux ./cmd/mcp-mux

# Windows
go build -o mcp-mux.exe ./cmd/mcp-mux

Place the binary somewhere on your PATH, or reference it by absolute path in .mcp.json.

2. Configure

Take any MCP server entry in .mcp.json and move the command into args[0], replacing command with mcp-mux:

Before:

{
  "mcpServers": {
    "engram": {
      "command": "uvx",
      "args": ["engram-mcp-server", "--db", "/data/engram.db"]
    }
  }
}

After:

{
  "mcpServers": {
    "engram": {
      "command": "mcp-mux",
      "args": ["uvx", "engram-mcp-server", "--db", "/data/engram.db"]
    }
  }
}

3. Verify

mcp-mux status

On the next CC session start, mcp-mux intercepts the stdio channel, connects to (or starts) the daemon, and proxies all MCP traffic transparently.

Sharing Modes

Mode Behavior Use When
shared (default) One upstream serves all sessions. Responses to initialize, tools/list, prompts/list, and resources/list are cached and replayed without a round-trip. Stateless servers: search, docs, LLM proxy.
isolated Each session gets its own upstream process. Per-session state: browser automation, SSH, editor buffers.
session-aware One upstream; sessions identified by injected _meta.muxSessionId. Stateful servers that can partition in-process state by session key.

Override mode for a specific server:

# Force isolation for one invocation
MCP_MUX_ISOLATED=1 mcp-mux uvx my-server

# CLI flag (equivalent)
mcp-mux --isolated uvx my-server

Modern R1 (MCP 2026-07-28)

Use R1 only when the host and upstream are both known to speak MCP 2026-07-28. It is an additive opt-in. Omitting --mcp-protocol=2026-07-28 preserves the legacy quick start and default behavior. R1 does not translate legacy and modern traffic.

Put --mcp-protocol=2026-07-28 before the upstream command:

mcp-mux --mcp-protocol=2026-07-28 my-modern-server --stdio

For example, configure the server in .mcp.json as follows:

{
  "mcpServers": {
    "modern-server": {
      "command": "mcp-mux",
      "args": [
        "--mcp-protocol=2026-07-28",
        "my-modern-server",
        "--stdio"
      ]
    }
  }
}

R1 forces isolation. It disables response caching, discovery templates, and replay. There is no automatic fallback to the legacy protocol or a shared mode. Server log notifications stay request-scoped and require the request to opt in with _meta.io.modelcontextprotocol/logLevel; mcp-mux does not broadcast them.

After an owner or transport loss, mcp-mux returns errors for in-flight modern requests. The host must issue fresh retries and a fresh subscriptions/listen; R1 does not replay requests or subscriptions.

To inspect the R1 policy, run mcp-mux status and look for these four fields:

  • protocol_era: "2026-07-28"
  • sharing_policy: "forced-isolated"
  • cache_policy: "off"
  • lifecycle_policy: "r1-quarantine"

R1 is separate from the historical --stateless flag. --stateless changes legacy server identity and sharing. It does not select an MCP protocol era or make R1 shareable.

To roll back an R1 configuration:

  1. Remove --mcp-protocol=2026-07-28 from the host configuration to stop new R1 admissions.
  2. Let current R1 owners drain, or remove them with mux_stop by the server_id shown in mcp-mux status. mux_stop is the existing control-plane tool for any owner. R1 adds no separate stop path. mcp-mux stop --drain-timeout 30s drains and stops the whole local daemon.
  3. Start a new host connection with the legacy configuration. Never downgrade a live R1 owner or replay old requests or subscriptions.

A refused, absent, or mismatched era confirmation means modern admission failed. Do not retry the same route as legacy.

Auto-Classification

When no explicit mode is set, mcp-mux classifies each server automatically using this priority order:

  1. x-mux capability (highest) — server declares x-mux.sharing in its initialize response. Authoritative; overrides all heuristics.
  2. Tool-name heuristics — tools with names matching browser, session, editor, navigate, page, tab, process, document, or snapshot patterns trigger isolation.
  3. Default — shared.
flowchart TD
  A[Server starts] --> B{x-mux capability\nin initialize response?}
  B -->|Yes| C[Use declared mode]
  B -->|No| D{Tool names match\nisolation patterns?}
  D -->|Yes| E[Isolated]
  D -->|No| F[Shared]
Loading

If your server is stateless but has tool names that match isolation patterns, add "x-mux": { "sharing": "shared" } to your initialize capabilities to fix the classification.

Response Caching

In shared mode, the owner intercepts and caches the first response for each of these methods:

  • initialize
  • tools/list
  • prompts/list
  • resources/list
  • resources/templates/list

Subsequent sessions receive the cached response immediately without a round-trip to the upstream. Cache entries are invalidated when the upstream sends the corresponding *_changed notification (notifications/tools/list_changed, notifications/prompts/list_changed, notifications/resources/list_changed).

For initialize, the cache is keyed on protocolVersion. A new client using a different protocol version bypasses the cache and goes to the upstream directly.

Cache-Only Startup and Lazy Materialization

The first time mcp-mux sees a command and security context, it starts the upstream and performs a synthetic initialize plus tools/list to learn the server's sharing mode and publish a reusable discovery template.

After that owner is removed, a compatible session can start as a cache-only owner:

  • cached initialize, tools/list, and any captured prompt/resource discovery responses are replayed without starting the upstream process;
  • notifications/initialized is suppressed while no upstream exists;
  • the first uncached request starts exactly one upstream generation and forwards the original request on the same open host transport after initialization;
  • mcp-mux status reports materialization_state: "CACHE_ONLY" with upstream_pid: 0 until demand arrives, then READY with the live PID.

Template reuse is fail-closed. A missing, stale, or incompatible template takes one bounded cold/eager path instead of replaying cached frames from the wrong context. Compatibility uses an exact SHA-256 identity of the effective security-relevant environment; isolated templates also require the exact canonical working directory. Raw environment values are not stored in the identity or exposed in status. Persistent owners and callers that explicitly require eager startup continue to materialize without waiting for a request.

For slow-starting servers (serena via uvx ~3s, tavily via npx ~5s), a compatible template lets later host startup complete from cache while deferring that cost until the server is actually used.

Daemon Mode

The daemon is enabled by default. It starts automatically when the first mcp-mux shim connects and no daemon is running.

Lifecycle:

  • Shim connects → daemon starts or is reused.
  • A non-persistent initialized shim with no requests or queued work parks its daemon IPC session after 10 minutes without host traffic. New host demand reconnects to the exact owner.
  • If no demand arrives for another 30 seconds, the stable launcher parks the engine process too. The next host frame starts the current active engine, replays the cached initialize handshake, and forwards that demand once.
  • When the last session disconnects, a disposable owner is eligible for the existing 30-second safety-gated zero-session cleanup. The general owner idle timeout remains 10 minutes; isolated owners use their shorter lifecycle rule.
  • Servers declaring x-mux.persistent: true do not suspend or reap while idle; they stay alive until explicitly stopped or until the daemon exits.
  • Daemon auto-exits after 5 minutes with no owners and no connected sessions.

MCPMUX_SHIM_IDLE_TIMEOUT and MCPMUX_SHIM_DORMANT_GRACE override the 10-minute and 30-second product-shim stages with Go duration strings. Zero or a negative value disables that stage; invalid values keep the default. These are separate from owner cleanup and daemon auto-exit settings.

Serena's web dashboard is configured separately from mux lifecycle. To prevent it opening automatically, pass --open-web-dashboard false to Serena's start-mcp-server command or set web_dashboard_open_on_launch: false in serena_config.yml; this leaves the dashboard active. Set web_dashboard: false only when the dashboard itself must be disabled; see the Serena dashboard documentation.

Standalone upstream launches through MCP_MUX_NO_DAEMON=1, MCP_MUX_DAEMON, or the direct-owner --daemon path are explicitly maintenance_unsupported in v0.31.0. Keep managed daemon admission enabled; direct execution is not a maintenance bypass.

Hold an upstream for executable replacement

v0.31.0 adds optional maintenance-aware managed hold, resume, and renew operations. Use an aware binary, daemon, and shim together. Publication and consumer handoff remain pending; older installed tags do not acquire maintenance support.

Read the exact server_id from local mcp-mux status. Flags follow the identifier:

mcp-mux hold <exact-server-id> --ttl 5m --drain-timeout 10s --json
mcp-mux renew <returned-hold-id> --ttl 5m --json
mcp-mux resume <returned-hold-id> --json

Hold fences new demand before draining already-forwarded requests. The default drain is 10 seconds; --drain-timeout 0s skips grace but still requires complete managed tree death. Positive grace uses the same T as TTL and never restarts on retirement retry. Replace the executable yourself only after a successful JSON result reports state: HELD, trees_retired: true, and a future expires_at. The result also contains hold_id, server_id, and drain_deadline.

State Meaning
HOLDING Admission is fenced; provisional timing precedes clocked drain/retirement. No replacement grant.
HELD Every scoped managed tree is proven dead; the lease is usable until expiry.
RETIREMENT_BLOCKED Tree death is unproven; replacement is unsafe and admission stays fenced.
RELEASED Durable release permits fresh demand.

TTL defaults to five minutes, must be positive, and cannot exceed one hour. Renewal uses the exact current unexpired hold ID and sets expiry from serialized renewal acceptance, without changing retirement state or reviving a released/expired lease. CLI TTL and drain durations must be whole milliseconds; sub-millisecond values are rejected rather than rounded down to zero-force retirement. Competing or stale identities cannot replace or clear a lease. Explicit resume and TTL recovery open admission only after proven retirement and durable release. Blocked retirement never clears because time elapsed.

A durable HOLDING seed has provisional timing and never grants replacement. After its first complete writer acknowledgment, sample T once and persist clocked HOLDING once. TTL/drain use T; that write, retirement, HELD persistence, and response consume the original window. Any acquisition write failure retains the seed fence; incomplete recovery is RETIREMENT_BLOCKED, with no expiry/resume.

With no positive caller timeout, neutral control allows 180s plus one drain for hold/restart_owner; CLI/MCP use this default. Explicit positive library budgets are honored; other commands retain 5s. This finite exchange allowance does not guarantee full-pin/storage completion. A timeout leaves outcome unknown: a durable lease/restart may remain. Inspect status, with no automatic retry/resume or stop fallback.

Aware shims keep the original host pipes open. Requests received while fenced return JSON-RPC -32005, message upstream held for update, and data.error_code: maintenance_held with the original numeric or string ID. Maintenance wins over cached success. Rejected requests and unfinished work ended by retirement are not replayed; notifications receive no invented ID. Fresh legacy demand after release can reach a replacement on the same pipes. Modern demand uses fresh same-era isolated admission with required per-request metadata, without legacy bootstrap, cache, or request/subscription restoration.

Scope is the selected owner's finite already-admitted context set, not a host-wide executable lock. Different CWDs, protocol eras, credentials/configuration contexts, engine namespaces, and unmanaged processes are outside that set unless already admitted explicitly. An unrelated process can still lock the file.

MCP mux_hold, mux_resume, and mux_renew use the same local daemon operation. mux_restart uses daemon-owned restart_owner with the original exact context and era, not an adapter's ambient credentials. Controlled restart, handoff, shutdown, downgrade, and idle daemon exit refuse terminally while a fence remains; launcher and library update helpers cannot substitute shutdown or a successor. After unplanned daemon loss, aware startup reloads held authority before admission. Incomplete or unreadable authority fails closed.

Durable authority is the mandatory schema-2 pair ledger.json and transaction.json, not either file alone. Pending, missing, or invalid pairs fail closed. A persistence error keeps live admission conservative but does not guarantee storage rollback. After a finalize error, recovery can accept only a matching COMMITTED certificate proving earlier acknowledged durable publication; the failed caller response remains an error. Successful release cannot resurrect the old lease. See the storage phase contract.

Controlled engine installation, launcher swap, layout/bootstrap mutation, and active-pointer updates use the existing daemon namespace file lock. Hold-ledger mutations use the same lock, so a hold cannot race past an activation check. daemon.CheckMaintenanceForActivation reads persisted authority without changing it. Status and pure startup inspection neither acquire nor write that lock and never proactively start a daemon. Activation checks also consult live aware status; an offline or old endpoint is allowed only when persisted authority is proven clear under the lock. This is namespace coordination, not a host-wide file lock or a separate updater lease.

Old daemons return maintenance_unsupported, never a stop/kill/restart fallback. An aware daemon fences old managed shims' starts, but those shims have no promised immediate-error or non-replay semantics. Arbitrary old binaries, foreign engines, manual active-pointer replacement, and direct standalone bypass are unsupported. Before downgrade, use the current aware binary to durably resume every exact retired lease or observe its safely committed expiry. Incomplete/blocked authority prevents downgrade and must be retained, even after TTL. Do not delete either authority member or perform PID cleanup to reopen admission.

Run the live Windows and Unix replacement proof before release. Its private primary-checkout scratch and actual overwrite evidence supplement focused regressions; a fixture build or unit test is not that proof.

Resilient Shim

Outside maintenance fences, legacy mcp-mux shims automatically reconnect when the daemon restarts. Modern R1 does not replay the legacy handshake below. This means:

  • mcp-mux upgrade switches the active versioned engine without dropping connections
  • mcp-mux stop --force triggers automatic reconnect within seconds
  • Daemon crashes are recovered transparently

During reconnect, the shim:

  1. Detects IPC connection loss (daemon shutdown)
  2. Drains orphaned in-flight requests — sends spec-compliant JSON-RPC error responses so CC sees explicit failures for pending requests instead of silence (silence on a pending request is what CC's stdio transport tears the connection down over)
  3. Tries the planned reconnect path first, including reconnect-token refresh
  4. Starts or waits for the replacement daemon when needed
  5. Replays cached initialize request to warm the replacement owner
  6. Sends notifications/tools/list_changed so the host can refresh discovery
  7. Flushes any still-valid buffered requests and resumes normal proxy

Reconnect is transport continuity, not request replay. A request already sent to the lost owner receives one explicit JSON-RPC error with its original id and is never sent to the successor. Only the cached initialize handshake is replayed to warm the replacement connection; host frames accepted afterward are forwarded once.

Reconnect timeout: 30 seconds. If reconnect is still unavailable after that window and a reconnect path exists, the shim does not exit. It enters degraded retry: new client requests receive JSON-RPC errors by their original ids, the parent stdio transport stays open, and normal proxying resumes on that same transport after backend recovery. The shim exits only when the MCP host closes stdin/stdout or when no reconnect path was configured.

Note on keepalives: Earlier versions emitted synthetic notifications/progress with a mux-reconnect progress token every 5 s as a keep-alive. That violated the MCP spec (progress tokens must reference a client-issued _meta.progressToken), and Claude Code tore down the stdio transport on the first unknown token — destroying the connection the shim was trying to preserve. The keep-alive was removed in muxcore v0.19.6; drainOrphanedInflight is the spec-compliant replacement.

Session Transport Layer

mcp-mux v0.4.0 introduces a session transport layer that replaces the old lastActiveSessionID heuristic with deterministic, per-session routing.

Token handshake

When CC spawns a shim, the daemon generates a cryptographic token tied to that spawn's working directory. The shim sends this token as the first line on the IPC connection:

CC → shim → [token\n] → Owner (SessionManager) → upstream

The Owner reads the token, looks up the corresponding Session.Cwd, and binds the IPC connection to that session. From this point the session identity is authoritative — no heuristics required.

Handshake enforcement (v0.9.10+). The Owner rejects IPC connections with an empty or unregistered token when daemon mode is active. Rejections are logged at owner level with the peer PID (no token value) and rate-limited to 10 entries per minute per owner with a suppressed-count summary. Pre-registered tokens are preserved on rejection, so a legitimate client that closes mid-handshake can reconnect without forcing the daemon to re-issue a new token. Tokens are 128-bit (16 random bytes from crypto/rand); entropy failure is fatal.

Deterministic callback routing

The SessionManager tracks inflight requests per session. When exactly one session has pending requests outstanding, response routing is deterministic without needing to inspect message content. This eliminates spurious mis-routing in high-concurrency scenarios.

roots/list forwarding

roots/list requests from the upstream are forwarded to the active CC session (the one with pending requests), so the server receives the real workspace roots for that session rather than a static fallback.

Security Model

mcp-mux is designed for a single-user local trust boundary: any process running as the same OS user is implicitly trusted. Two layered defenses protect against same-machine impersonation on shared Unix hosts:

Application-layer: handshake enforcement

The Owner acceptLoop rejects IPC connections with an empty or unregistered token (daemon mode). Combined with 128-bit crypto/rand tokens and single-use Bind semantics, this closes the only application-layer impersonation gap on the data socket.

OS-layer: 0600 socket permissions (Unix)

All Unix domain sockets created by ipc.Listen and the daemon control socket go through the muxcore/sockperm package, which applies syscall.Umask(0177) under a package-level mutex — the socket file lands with mode 0600 and is only accessible to the owner UID. On Windows, AF_UNIX sockets inherit the creating process's default DACL (owner + LocalSystem), so no umask equivalent is needed and the package is a documented no-op.

What mcp-mux does NOT protect against

  • Malware running under the same user account. A process with your UID can still connect to your 0600 control socket and issue its own spawn request to obtain a fresh pre-registered token. Treat the control socket as trusted to everything running as you.
  • Network-level adversaries. mcp-mux uses Unix sockets / Windows AF_UNIX only — there is no TCP listener. Remote attack surface is zero.
  • Upstream MCP servers themselves. mcp-mux is a transparent proxy; if an upstream server runs exec.Command on attacker-controlled input, mcp-mux doesn't rewrite or sanitize that.

Multi-user deployment

For shared-machine Unix hosts (multiple login users), mcp-mux v0.9.10 and later is safe for the cross-user boundary — the 0600 permission prevents a different user from connect()-ing, and the token handshake rejects same-user probe attempts that haven't received a pre-registered token from the daemon.

Commands

# Show all running upstream instances (PID, sessions, classification, cache state)
mcp-mux status

# Stop all running instances and the daemon
mcp-mux stop [--drain-timeout 30s] [--force]

# Versioned engine upgrade (see section below)
mcp-mux upgrade

# Start a detached daemon process (normally auto-started by shims)
mcp-mux daemon

# Run as control-plane MCP server (exposes mux_list / mux_stop / mux_restart tools)
mcp-mux serve

Versioned Engine Upgrade with Safe Restart

Upgrading the mcp-mux binary while sessions are active is safe and fast, but ordinary upgrades do not force live host transports through a daemon restart. The configured mcp-mux executable is a stable launcher and stdio anchor; the runtime code lives behind the versioned engine pointer.

This section describes the mcp-mux product updater. If you embed muxcore in another product, start from muxcore/README.md and choose that product's own launcher, engine path, staging name, and status/update surface instead of copying the mcp-mux.exe~ / mcp-mux.versions layout blindly.

# One-command upgrade with safe restart/defer semantics
go build -o mcp-mux.exe~ ./cmd/mcp-mux && mcp-mux upgrade --restart

What happens:

  1. The stable mcp-mux launcher is left in place; it is not renamed while live shims hold it.
  2. The pending mcp-mux.exe~ engine is copied or moved into mcp-mux.versions/<hash>/mcp-mux-engine.exe.
  3. mcp-mux.versions/active.txt is switched to the new engine path.
  4. If the daemon has live sessions, the daemon restart is deferred. Existing host stdio transports stay attached to their current daemon/engine boundary, and new shims use the new active engine pointer.
  5. If zero live sessions can be proven, --restart may perform graceful restart: the daemon serializes state snapshot (cached init/tools/prompts/resources responses, classification, session metadata), shuts down, and starts the successor daemon from the new engine.
  6. Current-generation shims reconnect to the successor, refresh reconnect tokens, and resume against restored owners with pre-populated caches.

The launcher indirection is intentional on Windows: live mcp-mux.exe shim and daemon processes keep the executable image locked, so a self-rename of the configured binary is not a reliable update primitive. Versioned engines avoid that lock; old processes keep running from their old engine path while new shims use the active engine pointer.

This has one bootstrap boundary: processes that were already launched by an older, pre-supervisor mcp-mux serve cannot be retrofitted in-place. Move the host configuration onto the stable launcher once; after that, ordinary engine updates preserve the host-facing stdio boundary by keeping the launcher stable and deferring daemon restart when live sessions are present.

Without --restart (active pointer switch only):

mcp-mux upgrade

The daemon keeps running with its current engine. New shim processes use the new active engine. The daemon updates on next natural restart.

When a same-protocol graceful restart is safe and actually runs, it preserves:

  • The upstream process tree — a same-v2 handoff retains it after successor adoption; a restart with live sessions is deferred instead
  • Cached MCP responses (init, tools, prompts, resources)
  • Server classification (shared/isolated/session-aware)
  • Session metadata (cwd, env)
  • Reconnect-token history, so live shims can refresh without fallback-spawning an owner during planned restart.

Only the daemon restarts — upstreams are reattached via FD passing (Unix SCM_RIGHTS, Windows DuplicateHandle). See the next section for the lifecycle contract.

Upstream Lifecycle — Transactional Restart and Full-Tree Cleanup

Handoff protocol v2 keeps an upstream process tree alive across a planned same-v2 daemon restart only after the successor adopts both stdio and the tree authority. Ordinary reconnect after an owner loss preserves the host transport but does not replay already-sent requests.

The contract

Trigger Pre-v0.21.0 Current contract
mcp-mux upgrade --restart with live sessions Upstream killed + respawned, in-flight requests dropped Daemon restart deferred; existing stdio transports remain on the current daemon while new shims use the new engine pointer
mcp-mux upgrade --restart with zero live sessions Upstream killed + respawned, in-flight requests dropped Same-v2 handoff retains the upstream tree; the first v1-to-v2 restart takes one snapshot-backed respawn
Daemon/owner loss Upstream lifecycle was leader-oriented Recovery is demand-driven; abandoned generations are cleaned as full trees and already-sent requests receive explicit errors instead of replay
mux_restart <sid> (operator-initiated) Hard kill Daemon-owned restart_owner preserves the exact launch context and era, with 30s drain by default. force: true skips drain only; maintenance-held targets refuse terminally without stop/spawn fallback.
Reaper idle-eviction Hard SIGKILL Soft-close: 30s stdin drain → SIGTERM only after timeout

How it works

*Unix (Linux, macOS, BSD):

  • Upstream spawns with Setpgid=true — the kernel places the child in its own process group.
  • Planned restart: old daemon opens a Unix domain socket, successor daemon connects with a 128-bit shared token, FDs (stdin, stdout, stderr) transfer via the SCM_RIGHTS ancillary control message.
  • Cleanup targets the process group, including descendants that outlive the leader or inherit its stdio. Planned v2 handoff retains the PGID authority until the successor's final adoption acknowledgment.

Windows:

  • Each upstream starts suspended and is placed in its own anonymous Job Object with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE before it can execute.
  • Planned restart: successor is spawned with a named-pipe address; handles are duplicated via DuplicateHandle with DUPLICATE_SAME_ACCESS, including the Job authority.
  • The predecessor retains its Job lease until the successor commits adoption; abort or final authority loss terminates the whole tree.

Handoff protocol

Old daemon → successor handshake is JSON-over-socket with a mandatory protocol_version: 2 field on every message:

Hello ──(token, source_pid)──>
       <──(protocol_version check, refs list)── Ready
FdTransfer ──(server_id, stdio + tree-authority metadata)──>
             <──(SCM_RIGHTS / DuplicateHandle)── AckTransfer (ok/aborted)
       ...repeat per upstream...
Done   ──(transferred, aborted lists)──>
       <──(accepted + aborted partition)── HandoffAck
  • Token auth (FR-11): constant-time compare, 128-bit random, 0600 file.
  • Per-upstream atomicity (FR-7): receipt does not detach the predecessor. Each tree commits only if the final acknowledgment accepts its server id; every other prepared tree aborts and is eligible for snapshot respawn.
  • 30s accept + total timeout on both sides.
  • Version skew (FR-3): negotiation happens before any owner detaches. A mismatched protocol_version, token mismatch, listener timeout, or handoff successor start failure aborts the restart, releases restart pins, and leaves the predecessor serving.

FR-8 degraded fallback

Snapshot fallback is authorized only after exact version/token Hello and owner detach. On a later receipt or final-ack failure, the daemon:

  1. proves the failed handoff successor has exited;
  2. aborts the prepared transfer, permanently closing the detached predecessor owner and terminating its process tree;
  3. rewrites the pinned snapshot from the retained logical generation metadata;
  4. pre-starts exactly one clean snapshot successor; and
  5. only then returns the post-response predecessor shutdown callback.

No-process/cache-only restarts likewise pre-start one snapshot successor before predecessor shutdown. A pre-detach failure logs handoff.abort and does not spawn fallback. A successful post-detach recovery logs handoff.fallback; failure to prove successor exit, rewrite the snapshot, or start the backup logs handoff.fallback_blocked and remains fail-closed.

The snapshot successor eagerly restores upstreams; drainOrphanedInflight returns JSON-RPC errors by original id to in-flight callers. It does not replay those requests.

Operator visibility

New counters in mux_list / HandleStatus:

Counter Meaning
handoff_attempted Total HandleGracefulRestart invocations that entered the handoff path
handoff_transferred Successfully handed-off upstreams across all handoffs
handoff_aborted Upstreams that fell back per-upstream (FR-7) while siblings succeeded
handoff_fallback Whole-handoff failures that took the FR-8 respawn path

Structured log markers: handoff.start, handoff.upstream.transferred, handoff.complete, handoff.fallback, handoff.receive.{start,complete,fail}.

Migration to v0.27.0

No consumer code change is required. The first restart from a v1 handoff binary to v0.27.0 rejects live transfer before detach and uses one bounded snapshot-backed respawn. Subsequent v2-to-v2 planned restarts retain stdio and tree authority transactionally. Rollback across the same v2/v1 boundary uses the bounded respawn again; do not force mixed-version live handoff.

Snapshot back-compat: v0.20.x OwnerSnapshot files load without errors; new fields (UpstreamPID, HandoffSocketPath, SpawnPgid) are omitempty and default to zero on old snapshots.

Known limitations

  • Protocol transition boundary: the first handoff v1 → v2 restart, and a rollback across that same boundary, reject live transfer before owner detach and take one bounded snapshot-backed shutdown-and-respawn path. Same-v2 restarts retain stdio and full-tree authority only after final successor adoption.
  • Per-upstream 30s transfer bound: upstreams that don't drain within 30s fall back to respawn for that entry only.
  • macOS launchd cross-parentage: verified via CI; spawns outside the mcp-mux process tree inherit correctly.
  • Windows process tree: each upstream is governed by a Job Object with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE; final authority loss intentionally terminates descendants, including those that outlive their leader.

Post-deploy verification

# Unix
scripts/verify-handoff.sh

# Windows
scripts\verify-handoff.ps1

The script spawns a test daemon, triggers upgrade --restart, asserts all upstream PIDs survive across the restart, and reports any dropped FDs.

muxcore library integration

Developers embedding muxcore in another Go MCP server should start with muxcore/README.md. The recommended integration path is engine.New + engine.Run: muxcore then owns daemon/client/proxy mode selection, token handshake, reconnect, session routing, snapshot restore, and graceful restart.

The lower-level handoff functions in muxcore/daemon remain public for custom daemon supervisors, but ordinary consumers should not build their own shim or connect directly to owner IPC paths. Doing so bypasses the daemon-minted session token path and will be rejected by daemon-managed owners.

See muxcore/README.md for the consumer checklist, required engine.Config fields, upgrade/launcher contract, guardrails, and low-level handoff API. The guide also states the muxcore API design rule: infer safe behavior in engine when possible, and fail early with an actionable error when the consumer wiring is ambiguous.

Reference

  • Spec: .agent/specs/upstream-survives-daemon-restart/spec.md
  • Engram: #109 (arc resolution), #130 (public API export for aimux-class consumers)

Configuration

All configuration is via environment variables. No config file is required.

Variable Default Description
MCP_MUX_NO_DAEMON 0 Legacy setting; 1 is unsupported and fails closed with maintenance_unsupported. Keep daemon-managed admission enabled; this is not a daemon bypass.
MCP_MUX_ISOLATED 0 Set to 1 to force isolated mode for this invocation
MCP_MUX_STATELESS 0 Set to 1 to ignore cwd in server identity hash (enables global deduplication)
MCPMUX_SHIM_IDLE_TIMEOUT 10m Safe host-idle period before a non-persistent shim parks its daemon IPC session; zero or negative disables
MCPMUX_SHIM_DORMANT_GRACE 30s Exact-owner reconnect window before a supervised launcher becomes dormant; zero or negative disables
MCPMUX_LAUNCHER_DORMANT_LEASE disabled Explicit opt-in to exit a dormant launcher after this additional no-demand lease. Use only with a host proven to relaunch after transport closure.
MCP_MUX_OWNER_IDLE 10m General owner idle timeout; overridden per owner by x-mux.idleTimeout
MCP_MUX_GRACE 10m Legacy alias used only when MCP_MUX_OWNER_IDLE is unset
MCP_MUX_IDLE_TIMEOUT 5m Daemon auto-exit after this period with no activity

Transport lifecycle and host boundary

MCP stdio has no standard logical-completion signal: the host owns termination by closing child stdin, while a server may send notifications or requests during an otherwise silent interval. mcp-mux therefore exits on host EOF but does not kill a silent live stdio transport by default. A capable launcher parks its engine after the shim idle/grace sequence and retains a small launcher stub for later demand. MCPMUX_LAUNCHER_DORMANT_LEASE bounds the complete disposable launcher/engine tree only as an explicit host-compatibility opt-in.

The installed stable launcher and active versioned engine are distinct binaries and may differ byte-for-byte. Private dormant frames require protocol-v2 target-bound bilateral attestation: before spawning the engine, the launcher binds a one-shot current-user local IPC endpoint, then binds the exact PID returned by the child start. The launcher accepts the fixed proof only from that client PID; the engine independently requires the endpoint server PID to be its OS direct parent and completes the fixed request/response exchange on that side channel, never on host stdio. The provider-derived version-store layout, active-engine pointer, and direct-parent executable path must also match the installed stable launcher. Custom or copied engine paths fail closed.

Forwarding the endpoint environment through a v0.27.0-or-older launcher does not transfer capability: the endpoint server remains an ancestor rather than the engine's direct parent. That running session stays fail-closed and receives no private dormant frames. The verified child may bootstrap the stable launcher for future invocations with the rollback-capable two-rename swap; restart only a host/session still running under the pre-v2 launcher before expecting launcher dormancy or a lease.

Customer-mode proof remains required for host relaunch behavior and live Windows executable swapping. Windows verifies the named-pipe server with GetNamedPipeServerProcessId, Linux uses SO_PEERCRED, and macOS uses LOCAL_PEERPID; the existing parent-image checks use the Windows process image, /proc/<ppid>/exe, and kern.procargs2 respectively. Unsupported platforms, including BSD targets without both proofs, fail closed and never emit private dormant frames.

The prepared current library target is muxcore/v0.31.0, which adds optional managed maintenance to the existing stable stdio supervisor and explicit native MCP 2026-07-28 route. After publication and Go proxy tag resolution, pin:

go get github.com/thebtf/mcp-mux/muxcore@v0.31.0

Ordinary legacy engine.New consumers require no source changes. Use control.SendMaintenance only when adopting maintenance. Select engine.Config.ProtocolPolicy explicitly for a known same-era modern route; legacy remains the zero-value default. Supervisor users should keep using supervisor.Run, supervisor.StartCommand, supervisor.StartWithFallback, supervisor.ProtocolV2, and the attestation package rather than copying the product adapter or private wire codec. Roll back to muxcore/v0.30.0 or a compatible prior binary only after exact retired leases are durably resumed or safely expired through the current aware version. Preserve incomplete/blocked authority and retire modern owners through quarantine, never live conversion.

The shared daemon is owned by the stable launcher rather than by any supervised engine generation. The launcher prepares it before starting a child; on Windows the daemon therefore remains outside the child's KillOnJobClose Job Object. A supervised child never spawns that daemon inside its own process tree and exits back to the stable launcher if the daemon must be recreated, preserving the host-facing stdio pipe.

Control Plane MCP Server

mcp-mux serve exposes an MCP server on stdio with management tools. Add it to .mcp.json like any other server:

{
  "mcpServers": {
    "mcp-mux": {
      "command": "mcp-mux",
      "args": ["serve"]
    }
  }
}

Tools:

Tool Description
mux_engines Lists opted-in native muxcore daemon engines registered on this host. Each descriptor is advisory and is verified by daemon status before being marked healthy. Stale or mismatched descriptors are labeled instead of mixed into owner lists. duplicate means more than one healthy descriptor advertises the same engine name; stale leftovers do not make a healthy daemon duplicate.
mux_prune_engines Dry-run by default. Lists or removes stale / invalid native muxcore registry descriptor files after the same verification used by mux_engines. This is registry garbage collection only: it never stops processes, owners, daemon control sockets, or live native muxcore products.
mux_list Returns running instances for the current project inside this mcp-mux daemon namespace (filtered by caller's cwd). Pass all: true to list this daemon's instances across all projects. Pass exact engine_name from mux_engines to query one verified native muxcore engine explicitly. Includes server ID, engine name, PID, downstream session count, pending requests, classification, and cache status. With verbose: true, includes classification source/reason and inflight request details when present.
mux_stop Gracefully drains and stops an instance by server_id. Use force: true for immediate kill. CR-001 scope is current mcp-mux daemon namespace only; it does not stop native registered engines.
mux_restart Uses daemon-owned restart_owner with the original exact launch context and era. Drain defaults to 30s; force: true skips drain only and never bypasses maintenance-held terminal refusal or enables stop/spawn fallback. Without an explicit target, resolves to the caller's session instance (e.g. mux_restart(name: "aimux") targets this project's mux-managed aimux, not a native engine). CR-001 scope is the current namespace only; cross-engine restart is a future opt-in feature.

Session-scoped control plane:

The control plane is session-aware. Each tool call is resolved in the context of the calling session's working directory:

  • mux_list — shows only servers owned by the current project by default. Use mux_list(all: true) for a full view across all projects in this mcp-mux daemon.
  • mux_engines — shows native muxcore products only when they explicitly opt into daemon registry advertisement.
  • mux_prune_engines — shows prune candidates with dry_run: true by default. Pass dry_run: false only after reviewing candidates; it removes stale / invalid registry descriptors, not daemon processes.
  • mux_list(engine_name: "aimux") — queries exactly one registered engine after verifying that the descriptor's control socket returns matching engine_name from daemon status.
  • mux_restart(name: "aimux") — resolves to the aimux instance started from this project's directory, not a same-named server from a different project.

This prevents accidental cross-project interference when multiple projects use the same server name.

The sessions count in mux_list is a downstream MCP client/shim count, not a count of visible terminal windows or top-level agent sessions. Some clients run hidden stdio app-server processes; each one may attach to mcp-mux as a separate downstream session.

Native muxcore products such as aimux or engram run under their own engine namespaces when they embed muxcore directly. They do not appear in default mux_list unless they were launched through the mcp-mux product daemon. If a native product opts into muxcore daemon registry advertisement, mux_engines can discover it and mux_list(engine_name: "...") can list that one engine's owners. Product-native health, sessions, upgrade, and restart surfaces remain authoritative unless that product later opts into explicit cross-engine management capabilities.

Prompts:

Prompt Description
mux-guide Full reference on architecture, classification, caching, and troubleshooting.
mux-status-summary Calls mux_list and returns a human-readable summary.

For MCP Server Authors

Declare your server's sharing preference in the initialize response capabilities:

{
  "protocolVersion": "2025-11-25",
  "capabilities": {
    "tools": {},
    "x-mux": {
      "sharing": "shared"
    }
  }
}

For stateless servers that don't depend on the client's working directory, add "stateless": true to enable global deduplication — one upstream instance regardless of which directory CC is opened from:

{ "x-mux": { "sharing": "shared", "stateless": true } }

For session-aware servers, mcp-mux injects into every request:

  • _meta.muxSessionId — unique session identifier (format: sess_ + 8 hex chars)
  • _meta.muxCwd — the CC session's project directory (for --project-from-cwd servers)
  • _meta.muxEnv — per-session environment variable diff (API keys, config paths)
{ "x-mux": { "sharing": "session-aware" } }

For servers that must stay alive across all session disconnects (e.g., expensive initialization, background indexing), declare persistence:

{ "x-mux": { "sharing": "shared", "persistent": true } }

Full protocol specification including implementation examples (TypeScript, Python, Go) and migration path: docs/mux-protocol.md.

Smoke Testing

mcp-mux includes a smoke test that validates mux-specific behavior with real upstream servers:

# Basic: verify serena works through mux
SMOKE_CWD=D:/Dev/my-project SMOKE_EXPECT=isolated \
  go run testdata/smoke_isolated.go uvx --from git+https://github.com/oraios/serena \
  serena start-mcp-server --project-from-cwd

# Isolation check: two projects get separate owners
SMOKE_CWD=D:/Dev/project-a SMOKE_CWD2=D:/Dev/project-b SMOKE_EXPECT=isolated \
  go run testdata/smoke_isolated.go uvx --from serena ...

# With tool call
SMOKE_CWD=D:/Dev/my-project SMOKE_TOOL=activate_project \
  go run testdata/smoke_isolated.go uvx --from serena ...

What it validates (mux behavior, not upstream correctness):

  • Spawn via daemon with proactive init
  • Classification matches expected mode
  • Session isolation: different cwds → different owners for isolated servers
  • Init response forwarded correctly through mux
  • Optional: tool call forwarded and response returned

Contributing

# Run tests
go test ./...

# Run vet
go vet ./...

# Build
go build ./cmd/mcp-mux

Pull requests are welcome. Please ensure go test ./... and go vet ./... pass before submitting. For significant changes, open an issue first to discuss the approach.

License

MIT

About

Transparent stdio multiplexer for MCP servers — share one upstream across multiple Claude Code sessions

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages