Skip to content

Commit 8990ee0

Browse files
chrisleekrclaude
andauthored
feat(orchestrator): add CLOSE_WAIT socket-spin watchdog for #264 signature (#266)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
1 parent 76e7810 commit 8990ee0

14 files changed

Lines changed: 2501 additions & 21 deletions

‎docs/operate/configuration.md‎

Lines changed: 26 additions & 21 deletions
Original file line numberDiff line numberDiff line change
@@ -74,26 +74,31 @@ Required whenever the orchestrator role is active.
7474

7575
## Orchestrator and daemon
7676

77-
| Variable | Default | Notes |
78-
| ------------------------------ | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
79-
| `WS_PORT` | `3002` | Orchestrator WebSocket listener. Must differ from `PORT`. |
80-
| `ORCHESTRATOR_URL` | _none_ | Presence flips the process to daemon mode. Use `wss://` in production; `ws://` emits a warning. |
81-
| `ORCHESTRATOR_PUBLIC_URL` | _none_ | Public WebSocket URL the spawner injects into ephemeral Pods. |
82-
| `DAEMON_AUTH_TOKEN` | _none_ | Shared secret for the daemon ⇄ orchestrator handshake. Required on both sides. Compared in constant time. |
83-
| `DAEMON_AUTH_TOKEN_PREVIOUS` | _none_ | Optional rotation overlap. Orchestrator accepts either the primary or this previous token; daemons always send the primary. See [`runbooks/daemon-fleet.md`](runbooks/daemon-fleet.md#rotating-daemon_auth_token). |
84-
| `HEARTBEAT_INTERVAL_MS` | `30000` | Daemon → orchestrator ping cadence. |
85-
| `HEARTBEAT_TIMEOUT_MS` | `90000` | Eviction threshold. Keep `≥ 2 × HEARTBEAT_INTERVAL_MS`. |
86-
| `FLEET_SNAPSHOT_INTERVAL_MS` | `30000` (clamp 10000-300000; `0` disables) | Cadence of the periodic `fleet.snapshot` gauge log (queue depth / daemon counts / free + busy slots). `0` disables it (inline-mode local dev). See [Fleet snapshot fields](observability.md#fleet-snapshot-fields). |
87-
| `STALE_EXECUTION_THRESHOLD_MS` | `3600000` | How long a `running` execution may sit before the watcher fails it. Set `≥ AGENT_TIMEOUT_MS`. |
88-
| `DAEMON_DRAIN_TIMEOUT_MS` | `300000` | Post-`SIGTERM` window to finish in-flight work. Raise to `≥ AGENT_TIMEOUT_MS` for zero mid-run kills. |
89-
| `JOB_MAX_RETRIES` | `3` | Retries for transient daemon dispatch failures. |
90-
| `OFFER_TIMEOUT_MS` | `5000` | How long the orchestrator waits for a daemon to claim an offer. |
91-
| `QUEUE_WORKER_BACKOFF_MAX_MS` | `5000` | Upper bound on the queue-worker's sleep when no local daemon can take a job. |
92-
| `LIVENESS_REAPER_INTERVAL_MS` | `30000` (min `20000`) | Cadence of the heartbeat-based reaper. |
93-
| `DAEMON_UPDATE_STRATEGY` | `exit` | `exit`, `pull`, or `notify`. Advisory hint reported in the update response. |
94-
| `DAEMON_UPDATE_DELAY_MS` | `0` | Delay before graceful shutdown after an update signal. |
95-
| `DAEMON_MEMORY_FLOOR_MB` | `512` | Minimum free memory the orchestrator requires before dispatching. |
96-
| `DAEMON_DISK_FLOOR_MB` | `1024` | Minimum free disk the orchestrator requires before dispatching. |
77+
| Variable | Default | Notes |
78+
| --------------------------------- | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
79+
| `WS_PORT` | `3002` | Orchestrator WebSocket listener. Must differ from `PORT`. |
80+
| `ORCHESTRATOR_URL` | _none_ | Presence flips the process to daemon mode. Use `wss://` in production; `ws://` emits a warning. |
81+
| `ORCHESTRATOR_PUBLIC_URL` | _none_ | Public WebSocket URL the spawner injects into ephemeral Pods. |
82+
| `DAEMON_AUTH_TOKEN` | _none_ | Shared secret for the daemon ⇄ orchestrator handshake. Required on both sides. Compared in constant time. |
83+
| `DAEMON_AUTH_TOKEN_PREVIOUS` | _none_ | Optional rotation overlap. Orchestrator accepts either the primary or this previous token; daemons always send the primary. See [`runbooks/daemon-fleet.md`](runbooks/daemon-fleet.md#rotating-daemon_auth_token). |
84+
| `HEARTBEAT_INTERVAL_MS` | `30000` | Daemon → orchestrator ping cadence. |
85+
| `HEARTBEAT_TIMEOUT_MS` | `90000` | Eviction threshold. Keep `≥ 2 × HEARTBEAT_INTERVAL_MS`. |
86+
| `FLEET_SNAPSHOT_INTERVAL_MS` | `30000` (clamp 10000-300000; `0` disables) | Cadence of the periodic `fleet.snapshot` gauge log (queue depth / daemon counts / free + busy slots). `0` disables it (inline-mode local dev). See [Fleet snapshot fields](observability.md#fleet-snapshot-fields). |
87+
| `SOCKET_HEALTH_INTERVAL_MS` | `30000` (clamp 5000-300000; `0` disables) | Cadence of the CLOSE_WAIT socket-spin watchdog (issue #265). `0` disables it (e.g. no procfs in local dev). Does not fix #264, it detects and structurally logs the signature. See [Socket health watchdog events](observability.md#socket-health-watchdog-events). |
88+
| `SOCKET_HEALTH_LEAK_SAMPLES` | `3` (clamp 2-100) | Consecutive samples a CLOSE_WAIT socket must survive before it is logged as a leak. Lower is noisier; the floor of 2 stops a transient socket from being flagged. |
89+
| `SOCKET_HEALTH_SELF_HEAL_SAMPLES` | `10` (clamp 2-1000) | Consecutive samples a leak must persist, alongside a pinned core, before it is treated as a spin. |
90+
| `SOCKET_HEALTH_CPU_PERCENT` | `90` (clamp 50-100) | CPU floor for a spin, as a percentage of one core. CPU alone is never sufficient: a 13.5s `scheduler.scan` legitimately burns a core. Only persistent CLOSE_WAIT plus this floor escalates to a spin. |
91+
| `SOCKET_HEALTH_SELF_HEAL_ENABLED` | `false` | When `true`, a suspected spin exits the process with code `75` (EX_TEMPFAIL) so k8s restarts the pod and bounds the burn. The distinct code lets `lastState.terminated.exitCode` tell a self-heal from a real crash. |
92+
| `STALE_EXECUTION_THRESHOLD_MS` | `3600000` | How long a `running` execution may sit before the watcher fails it. Set `≥ AGENT_TIMEOUT_MS`. |
93+
| `DAEMON_DRAIN_TIMEOUT_MS` | `300000` | Post-`SIGTERM` window to finish in-flight work. Raise to `≥ AGENT_TIMEOUT_MS` for zero mid-run kills. |
94+
| `JOB_MAX_RETRIES` | `3` | Retries for transient daemon dispatch failures. |
95+
| `OFFER_TIMEOUT_MS` | `5000` | How long the orchestrator waits for a daemon to claim an offer. |
96+
| `QUEUE_WORKER_BACKOFF_MAX_MS` | `5000` | Upper bound on the queue-worker's sleep when no local daemon can take a job. |
97+
| `LIVENESS_REAPER_INTERVAL_MS` | `30000` (min `20000`) | Cadence of the heartbeat-based reaper. |
98+
| `DAEMON_UPDATE_STRATEGY` | `exit` | `exit`, `pull`, or `notify`. Advisory hint reported in the update response. |
99+
| `DAEMON_UPDATE_DELAY_MS` | `0` | Delay before graceful shutdown after an update signal. |
100+
| `DAEMON_MEMORY_FLOOR_MB` | `512` | Minimum free memory the orchestrator requires before dispatching. |
101+
| `DAEMON_DISK_FLOOR_MB` | `1024` | Minimum free disk the orchestrator requires before dispatching. |
97102

98103
## Ephemeral daemons
99104

@@ -204,7 +209,7 @@ before committing:
204209

205210
## Prompt cache layout
206211

207-
Selects the system/user prompt split the agent executor passes to the Claude Agent SDK. See `src/config.ts:582#promptCacheLayout` for the Zod definition and `src/core/executor.ts:208` for the runtime guard.
212+
Selects the system/user prompt split the agent executor passes to the Claude Agent SDK. See `src/config.ts:604#promptCacheLayout` for the Zod definition and `src/core/executor.ts:208` for the runtime guard.
208213

209214
| Variable | Default | Notes |
210215
| --------------------- | -------- | ------------------------------------------------------------------------------------------------------------ |

0 commit comments

Comments
 (0)