Skip to content

Latest commit

 

History

History
494 lines (342 loc) · 28.2 KB

File metadata and controls

494 lines (342 loc) · 28.2 KB

Troubleshooting

Current reference. Symptoms → likely cause → supported recovery. Historical browser sidecar and manual deploy instructions are intentionally not used here.

mso update says local main diverged from origin/main

This is a Git safety stop, not a systemd failure. MSO refuses to overwrite local commits. Use:

mso update reconcile
mso update

reconcile requires a clean main, fetches origin/main, preserves the previous local HEAD on a timestamped rescue/mso-update-* branch, then makes local main exactly match origin/main. It never deletes the preserved commit. On a container/Codespaces host without systemd, the subsequent update uses the loopback fallback runtime; systemd is optional.

If the local web/PTY runtime is down after a container restart, start the loopback runtime directly:

mso web --local --print

Then verify http://127.0.0.1:4005/api/health (or the configured loopback port). Public HTTPS is a separate layer: Cloudflare is optional; Tailscale Serve, a reverse proxy, or another externally managed HTTPS endpoint can be used instead.

Login & sessions

Login returns not_configured (HTTP 500)

OS_SESSION_SECRET is missing/too short or OS_LOGIN_PASSWORD is invalid. MSO fails closed. Fix .env.local, then restart through your normal service/update path.

"Too many attempts, try again later" (HTTP 429)

The per-IP login limiter tripped. Wait for the window to pass. Behind a reverse proxy, forward the real client IP consistently so every user does not share the proxy address.

Correct password but "pending approval"

Expected for a new browser. Approve the shown device id from an already-approved browser (Settings → Account → Devices) or from the server using the approval script. Pair and re-check from the same browser origin. Device identity is browser-origin scoped, so approving an ID from http://server-ip:4005 and then switching to https://mso.example.com creates a different ID.

Device was approved but login still does not stick

Do not use plain HTTP on a non-loopback hostname/IP. MSO's session cookie is always Secure; a browser can reach http://server-ip:4005 and submit the password, but it cannot retain the session cookie there, so the next authenticated probe stays signed out. Use HTTPS, or tunnel the service and open http://localhost:4005, then approve the device ID shown on that final origin and click Check again. mso doctor now calls out this transport mismatch; mso doctor --fix repairs safe local issues such as a pending CLI device/session, but deliberately does not change DNS, TLS, firewall rules or public exposure. If you approved the wrong-origin browser ID, revoke it only after the final-origin device is working.

mso onboard approved the CLI device, then says port 4005 is unreachable

Device approval and runtime liveness are separate. Current MSO does not tell you to approve the same device again for a connection failure: on loopback, onboarding asks the gateway runtime helper to verify the MSO health contract and, on WSL/no-service installs, start the already-built production runtime on loopback. If the build is missing or stale, run mso update, then mso web, then resume mso onboard. Running mso device approve <id> again with the same role is idempotent; changing an existing device role still requires the explicit mso device role command.

How do I update when the web UI / port 4005 is down?

Run mso update. The CLI updater reads/fetches origin/main directly and does not need the MSO API. On WSL without an active service it verifies and builds the clean updated checkout. Every gateway-owned fallback runtime for that canonical checkout is inventoried and quiesced before .next changes, then restored afterward while active tunnel identities are preserved. Checkout-scoped private deployment receipts/restart markers survive partial failures, so rerunning the same mso update retries dependency/build/restart work even when Git HEAD already equals origin/main. Offline transactions are serialized; do not remove the owner-only update-state directory just to bypass a pending recovery. Recovery intent is written before a gateway-owned runtime is quiesced, so an interrupted state update remains safely retryable. Use mso update status for source/deployment state and mso update log for the service-updater transcript. When mso.service is active, status also compares the commit baked into the live loopback /api/health response with source HEAD; source equality alone is not treated as proof that deployment finished.

mso service start or restart while mso web is serving locally

MSO takes the checkout runtime exclusion exclusively before service handoff. If the unit is inactive and its loopback port is occupied by a gateway-owned fallback, MSO quiesces only that recorded process, preserves any public tunnel, then starts the system service on the freed port. Exit status alone is not readiness: MSO waits for the loopback /api/health identity before clearing recovery state; if readiness fails, the unhealthy service is stopped and the fallback is restored. A gateway-owned fallback that exited naturally is also reconciled into durable restore intent on the next update/deploy. An unowned/manual responder is never killed automatically.

mso update says local main is ahead/diverged from origin/main

This is an authority refusal, not an updater failure. Normal update may deploy only the fetched origin/main release. Push the intended commit to origin/main, or reconcile/reset the checkout, then rerun mso update. MSO will not silently build a clean but unpushed local commit. If Git source is intentionally already correct and only the production build needs repair, mso update --rebuild rebuilds that selected clean checkout without fetching or rewriting Git history.

Re-running the installer while MSO is already running

Supported reruns now use the same checkout-wide mutation boundary as self-update before dependency install or .next build. A same-checkout service is quiesced and refreshed as part of the installer lifecycle; gateway-owned WSL/no-systemd fallbacks are inventoried, stopped and restored afterward with their original env-file identities. A service from another checkout, an unowned loopback runtime, or --no-service against an active same-checkout service is refused before mutation.

mso update says the active service belongs to another checkout

This is a safety refusal. A machine may contain multiple MSO clones, but an active mso.service has one canonical WorkingDirectory. Run the update from that checkout (normally the directory targeted by the installer) instead of letting a secondary clone rebuild itself and restart someone else's unit.

mso update says the selected loopback runtime is not safely update-owned

MSO found a healthy loopback responder while mso.service is inactive, but that process is not recorded as a gateway-owned fallback runtime. This is commonly a manual bun run start/next start. Stop that manual runtime first, then rerun mso update. MSO refuses here before dependency or .next mutation rather than rebuilding underneath a process that is actively serving the old build.

mso update says a gateway-owned runtime predates env identity

The fallback was created by an older MSO version that did not record which env file launched it. MSO refuses before stopping that runtime because guessing .env.local could restart it with different auth, provider or public-origin configuration. Re-run the same local launcher once with the env file that runtime uses (for example mso --env /same/private.env web), then retry the update. This rewrites only the owner-only gateway lifecycle state with the env file's canonical path/device/inode; env contents are never copied into state. Do not delete the private restore/update state to bypass this check.

mso gateway start says an offline update is mutating this checkout

This is an intentional safety exclusion. Every in-place mso update (service-active or offline) holds the checkout-wide runtime lock while .next can change, so mso web, onboarding fallback, gateway runtime recovery, and mso service start/restart cannot start Next from a partially-mutated build. Let the update finish, then rerun the same gateway/web/service command. mso deploy uses the same runtime-quiesce lifecycle and requires the active service to belong to the checkout invoking the command.

mso gateway status says external-running

This is a healthy, intentional state when OS_PUBLIC_ORIGIN reaches the exact same MSO health identity as the selected local runtime but MSO has no owned provider-process record. A reverse proxy, Tailscale, Kubernetes ingress, PaaS router, s6-managed tunnel, or another external supervisor may own that lifecycle. MSO inspects the route but will not stop, restart, or automatically adopt the external process. Use mso gateway status --json for the separate ownership/provider/process/local/public dimensions. See Gateway architecture.

mso gateway status says degraded

A running process is not sufficient proof of public readiness. Check publicHealth, localHealth, processHealth, and ownership in mso gateway status --json. identity-mismatch means the public health endpoint is a valid MSO deployment but not the selected local instance; fix routing rather than restarting blindly. Doctor/status do not mutate providers. An external gateway is report-only; managed recovery remains limited to processes whose ownership identity MSO can prove.

Public MSO works, but mso gateway doctor reports an unsafe local bind

These are independent dimensions. A working public reverse proxy/tunnel does not make 0.0.0.0:<port> safe. The normal deployment keeps the raw MSO app on loopback/private scope and publishes it through the chosen HTTPS provider. Fix the runtime/service bind separately; do not suppress the doctor finding merely because public health is green.

mso gateway start cannot install or verify cloudflared

Current MSO installs a reviewed cloudflared release automatically into ~/.mso/tools on first mso gateway start. The release URL and SHA-256 for supported Linux architectures are pinned in security/gateway-artifacts.env; the cached binary is re-hashed before reuse and auto-update is disabled. Run mso gateway install to retry the dependency step by itself. If outbound GitHub release downloads are blocked, fix that network policy or set MSO_GATEWAY_CLOUDFLARED to an explicit locally reviewed executable. Set MSO_GATEWAY_NO_AUTO_INSTALL=1 when policy requires manual provisioning. A brand-new Quick Tunnel hostname may also take several seconds to become reachable. mso gateway start waits up to 60 seconds by default and still requires the exact MSO health/runtime-instance contract; this is readiness tolerance, not a weaker health check. A local resolver can cache the initial NXDOMAIN longer than the record creation itself; temporary mode therefore has a Cloudflare DoH fallback that still verifies HTTPS for the generated hostname and the exact runtime nonce.

mso gateway misdetects local health when my shell uses an HTTP proxy

Current MSO explicitly bypasses configured HTTP/HTTPS proxies for the already-validated loopback /api/health probe. Public HTTPS readiness still follows normal proxy policy. If mso gateway doctor still cannot verify the local runtime, inspect the local bind/build rather than adding the public Quick Tunnel hostname to NO_PROXY.

Temporary gateway opens, but Terminal does not stream

Cloudflare Quick Tunnels are the no-custom-domain preview mode and do not support Server-Sent Events. MSO Terminal output is an SSE stream, so use a named Cloudflare Tunnel/stable HTTPS origin for full Terminal behavior. mso gateway status labels a Quick Tunnel as temporary.

mso gateway stop leaves my manually-started MSO runtime running

Expected. The gateway only stops a Next runtime when it launched that exact loopback process and recorded it as gateway-owned. An existing systemd/manual runtime is outside the gateway lifecycle. This prevents stop from terminating an unrelated or pre-existing process. If an update/installer is in flight, user-facing mso gateway stop waits for that checkout transaction to finish restoring its recorded fallbacks, then applies the stop under the gateway lifecycle lock; the updater cannot restart it afterward.

Login returns success but the browser is logged out immediately

The session cookie is Secure. Plain HTTP on a normal IP/hostname causes browsers to drop it. Use HTTPS/Tailscale Serve or access through a localhost SSH tunnel.

Existing sessions suddenly died

A changed OS_SESSION_SECRET invalidates all signed sessions. Confirm deployment automation is not regenerating it.

Deploy & build

UI is unstyled or JS/CSS chunks 404 after deploy

The running next start process and on-disk .next tree do not match. Do not keep mutating the live .next tree manually. After any active updater/finalizer has finished, run the supported recovery rebuild:

mso update --rebuild

Then verify /api/health and run the post-deploy smoke check. Developer changes should be released with bun run ship "<conventional commit>", which performs the verified build/restart sequence.

Build runs out of memory

A Next production build can need multiple GiB. Add swap or raise the Node heap for the build process, then run the supported build/release path again. Do not restart a healthy old service after a failed build.

"Another next build process is already running"

First confirm a release/self-update is not actually active. Multiple simultaneous builds against the production checkout are unsafe. If no process owns the build and only a stale lock remains, remove the stale lock, then use the normal update/release command.

A new route still 404s after deployment

Confirm Git HEAD, origin/main, /api/health build id and the deployment log all refer to the expected release. If the deployment is correct but the build tree is inconsistent, use mso update --rebuild rather than hand-restarting around a partial build.

Update button says a newer version exists forever

Check Settings → Account → About and ~/.mso/self-update.log. A successful self-update ends with UPDATE OK. Also verify only one production process is serving the public origin. If the browser cached an old service worker, unregister it once and hard reload after the server is confirmed healthy.

Files

"Folder is outside the writable area"

The target is outside OS_FS_WRITE_ROOTS. Widen the write jail deliberately in .env.local if needed; do not use the sensitive-path escape hatch as a convenience.

Folder tree cannot see an expected path

Reads are bounded by OS_FS_READ_ROOTS plus the process user's Unix permissions. Credential paths remain hidden even when they are physically beneath an allowed root.

"Access to credential files is blocked"

Expected. MSO blocks its own private state, .env* and sensitive-home paths through the normal file API. Edit such material through a trusted host-admin channel rather than teaching the web file manager to read it.

Terminal / Code integrated terminal

Terminal says too many sessions

MSO caps concurrent live PTYs. Close unused tabs. Normal page/tab navigation sends a close, and stale detached shells can be reclaimed after a grace period when capacity is exhausted; attached active terminals are not reclaimed to make room.

Terminal disappears after server restart

PTYs are processes, not persisted sessions. A server restart terminates them. The client can open a fresh shell; command history/persistent work should live in the shell/tool itself (e.g. files, tmux when intentionally used), not PTY memory.

Code editor Terminal button works but starts in the wrong directory

Integrated Code Terminal uses the editor/project working directory passed as cwd. Confirm the opened file/project path is inside an allowed writable root; otherwise host cwd resolution falls back/refuses according to the host policy.

MSO Agent terminal

MSO Agent shows an Error section after HTTP/API failure

MSO 1.12 treats recoverable terminal API/transport failures as an interaction state instead of silently ending the conversation. The Error section shows only a bounded/redacted summary and reports the mutation outcome as not_started, completed, or uncertain. The durable session, correlation rows, permission mode and an active composer draft are preserved.

  • not_started — no mutation in that turn was dispatched; fix the underlying API/runtime issue and continue.
  • completed — the mutation finished before the later failure; do not repeat it merely because the follow-up model/API step failed.
  • uncertain — a write/exec request may have reached the server before the transport failed; inspect the target state first, then retry only if verification proves the mutation did not happen. MSO never auto-retries this state.

If the service itself is unhealthy, use mso doctor, mso service status, or the relevant read/status tool before attempting another mutation. Never paste raw request bodies, cookies, tokens, or hidden transcript data into a retry.

Permission/details interaction was interrupted

In ask mode the default approval view is one compact Approval needed: <tool> — <action> line. Enter opens redacted-safe exact-call details and the canonical digest; a separate explicit allow or deny decision is still required. If the approval interaction fails before that decision, MSO keeps pending approval metadata but does not infer approval. Continue/re-open the action and make an explicit decision.

Camoufox Browser

Browser says Camoufox is not installed

The current Browser app uses the camoufox-vnc.service user unit plus scripts/camoufox-vnc-service. The retired Playwright browser daemon and OS_BROWSER_URL/OS_BROWSER_SECRET are not part of current MSO.

Check the prerequisites documented in docs/INSTALL.md: Camoufox, a headless X server, x11vnc, noVNC/websockify, the VNC password file and the user systemd runtime/linger setup.

Browser is installed but Off

That is the expected idle state. The user unit is intentionally not enabled at boot and has a finite lease. The human-facing Browser app may start it from Browser/Settings when needed. Agent website audits should normally use the verified camoufox-browse automation skill with a disposable profile instead, so the persistent logged-in VNC session can remain off.

Browser reports secure embedding is unavailable

The Browser requires a reserved split-origin host. Configure NEXT_PUBLIC_MANAGED_APP_HOST_TEMPLATE and OS_SESSION_COOKIE_DOMAIN, then provision DNS/TLS for the resolved Camoufox host (for example camoufox.mso.example.com). Do not restore the old same-origin /camoufox-vnc/* route; it is a deliberate 404 security boundary.

noVNC loads but the viewer reports missing metadata/assets

MSO's launcher builds a private runtime noVNC webroot that symlinks the distribution assets and provides the package metadata Debian omits. Restart the Camoufox session so the runtime webroot is rebuilt from the installed noVNC package.

Camoufox starts with no saved logins

Verify CAMOUFOX_PROFILE points at the persistent profile under the user's local share tree, not a cache directory. Stop the session before restoring the profile/session backup. Treat the profile as account credentials: it can contain live cookies.

Browser is slow on mobile

The app is a real remote Firefox canvas, not a responsive website renderer. MSO's container is responsive, but the remote browser itself still has a desktop viewport. Use landscape or zoom/pan as appropriate; do not "fix" it by exposing VNC credentials to the client model.

Managed applications

A managed app is reported "not installed" but you know it exists

Look for the diagnostic that says MSO cannot reach the owner systemd user bus. Detection fails closed for install/restore because rerunning an installer or restoring over a live unknown service is unsafe. Fix the user bus/linger/service environment first.

Installed app is stopped and shows no dashboard

Expected. MSO mounts the vendor dashboard only when there is a live upstream. Start the app; management/log/update actions remain under Details.

9Router is healthy, the old hostname works, but a new 9router.mso... host does not

Verify the loopback dashboard first (curl http://127.0.0.1:20128/api/version). The *.mso... route belongs to optional split-origin embedding and has independent DNS/TLS, OS_SESSION_COOKIE_DOMAIN, build-time host-template and re-login requirements. A 401 from that host does not mean the 9Router container is down.

App is healthy but dashboard is not embedded

Embedded dashboards are opt-in. Confirm NEXT_PUBLIC_MANAGED_APP_HOST_TEMPLATE and OS_SESSION_COOKIE_DOMAIN, DNS/TLS for the explicit app hosts, and sign in again after changing the cookie domain. With those variables unset, no vendor iframe is served by design. A direct 9Router public-IP UI exists only when NINE_ROUTER_EXPOSE_PUBLIC=1 was deliberately configured.

Update fails before invoking the upstream updater

The mandatory MSO pre-update state snapshot failed. Fix that first; the update is designed to abort before changing the app if it cannot create its recovery point.

Restore is refused

The app must be stopped and MSO must be able to prove that. The backup manifest/source and current state directory must match, and symlink collisions are rejected before writes. Follow the specific refusal instead of bypassing the guard.

Alfa / model providers

Alfa says no API key/provider configured

Configure a BYOK provider in Settings → AI, set the corresponding server environment key, or intentionally connect the optional openai-codex provider. These are model credentials, not MCP credentials.

Custom provider URL is rejected

MSO SSRF-checks custom base URLs and refuses private/link-local/unsafe destinations according to its custom-provider policy. Use a legitimate endpoint reachable under that policy.

OpenAI Codex sign-in and ChatGPT MCP are being confused

They are different flows. openai-codex under Settings → AI is an Alfa inference provider. The ChatGPT custom MCP app authorizes ChatGPT to call MSO tools through MSO's /oauth/*. See docs/MODELS-INTEGRATION.md and docs/CHATGPT-PLUGIN.md.

MCP / ChatGPT custom app

Where do I start?

Use Settings → MCP → Connect an AI client. It now derives the MCP URL from MSO's deployment-owned public origin, probes the MCP endpoint plus OAuth challenge/resource/server metadata, and shows numbered setup steps for ChatGPT, Codex, Claude Code, Cursor, Gemini CLI, VS Code, and a generic Streamable HTTP client. Advanced OAuth fields stay collapsed unless a client actually asks for them.

If Settings says the origin is local-only (127.0.0.1, localhost, or ::1), do not paste that URL into a cloud client. Keep MSO loopback-only and use mso gateway start for a temporary HTTPS endpoint, or configure a stable HTTPS tunnel/reverse proxy and set the same origin with mso gateway domain set https://…. Reopen Settings through that public origin before copying the client URL.

Settings → MCP Test connection is red

The test is not a generic ping. All four public contracts must agree on the same deployment: GET /mcp, the unauthenticated Bearer/OAuth challenge from POST /mcp, RFC 9728 protected resource metadata, and RFC 8414 authorization-server metadata. Streamable HTTP browser requests also have their Origin validated before authentication to block DNS-rebinding access to a local listener. Expand Advanced OAuth settings and compare the URLs if one badge is red. A proxy/domain mismatch is more likely than a token problem because this probe intentionally runs before OAuth authorization.

/mcp or OAuth discovery returns 404

OS_MCP_ENABLED=1 is not active in the running service (or demo mode forced MCP off). After changing config, use the supported rebuild/update path and recheck GET /mcp.

ChatGPT says the MCP server does not support OAuth

Check both well-known endpoints from the same public origin ChatGPT reaches. A 404 normally means MCP is disabled in the running MSO process.

OAuth opens but cannot complete

The consent page is a normal MSO browser page. Use an already-approved device and active MSO session, review the requested read/write/exec tier, then Allow.

ChatGPT still shows old tools after MSO changed

Compare Settings → MCP toolset version/hash/count, refresh/recreate the ChatGPT custom MCP app, and run Scan Tools. "Mark ChatGPT refreshed" in MSO is only an acknowledgement; it does not refresh ChatGPT remotely.

Tool exists but returns scope denied

The OAuth bearer was granted below that tool's tier. Reauthorize intentionally at the needed scope rather than raising OS_MCP_MAX_SCOPE blindly.

fs_upload_file fails

The bridge accepts a current ChatGPT-provided file reference, up to 20 MiB, and writes only inside OS_FS_WRITE_ROOTS. Temporary OpenAI download URLs expire and are host/type/redirect validated. Reattach/regenerate the file instead of supplying an arbitrary public URL.

Project/skill seems missing

Check the scan report. Project and skill enumeration are bounded; if truncated:true, continue with the returned cursor instead of concluding absence.

Where to look next

  • release/update: docs/DEVELOPMENT.md, docs/INSTALL.md
  • ChatGPT connector: docs/CHATGPT-PLUGIN.md, docs/MCP.md
  • managed apps: docs/MANAGED-APPS.md
  • Camoufox: claude-skills/mso-camoufox/SKILL.md
  • security: SECURITY.md

Camoufox is running but its viewer is blank

Check each boundary separately. mso api GET /api/v1/camoufox/service reports the user service and loopback viewer; mso api GET '/api/v1/camoufox/service?probe=viewer' checks public HTTPS transport through the authenticated API. The second operation never forwards session cookies or VNC passwords and does not authenticate a VNC client. A TLS failure must be fixed at the exact viewer hostname, certificate and reverse-proxy route; repeatedly starting an already running browser cannot fix it. Keep the viewer on its isolated origin, retain live Operator/Owner authorization, and never expose a bare noVNC listener or disable certificate verification as a workaround.

A certificate for *.example.com covers camoufox.example.com but does not cover camoufox.mso.example.com. MSO therefore derives the safe sibling outside a configured mso.example.com cookie namespace when possible. If that hostname is not routed in your deployment, set CAMOUFOX_VIEWER_ORIGIN to an isolated HTTPS origin with valid DNS/TLS; do not move it back inside the cockpit cookie Domain. Check edge coverage independently from the origin certificate. For Cloudflare full-setup Universal SSL, see its hostname coverage limits. DNS-only routing requires a publicly trusted origin certificate and the intended origin route; changing the record also changes which edge protections apply.

The browser session may stop at its configured service time limit. A failed transport probe is not a stopped process. Verify actual navigation independently using an isolated profile before attributing a saved-profile failure to the browser binary. Preserve saved logins and profile data; a passing clean-profile test is not a repaired saved profile.

For GitHub connection diagnosis, compare the returned authenticated login with the intended account, then verify permissions on the exact repository. A named MSO profile, a successful /user response, and permission to merge a repository are three different facts.

Private or split-DNS Camoufox viewers

A configured viewer hostname may resolve exclusively to RFC1918, carrier-grade NAT, or unique-local IPv6 addresses. MSO does not send a server-side request to these addresses: diagnostics report client-only and reachable: false. The browser may then load the configured, separate viewer origin and verify its own HTTPS/authentication. This is not a server reachability pass. Loopback, metadata, link-local and mixed DNS answers remain denied; there is no unauthenticated proxy or same-origin fallback.