Skip to content

Security: NVitlam/agent-deck

Security

SECURITY.md

Security posture

Agent Deck reads what Claude Code, OpenCode and Codex leave on disk, and renders it. It never writes to any of them, never launches or wraps any of them, and never sends anything anywhere. This document records what is enforced, how it is enforced, and how that enforcement was measured — not intentions.

Scope note: this file carries no vulnerability-disclosure address and no support commitment, because neither exists yet. What it does carry is every version window, every file and table that is read, and every one that is deliberately never opened.


1. The architecture is the guarantee

Three kinds of source — the third optional — and none of them is a network client.

source where it comes from what it answers
hooks a hook snippet you paste yourself POSTs to a loopback HTTP listener — Claude Code's and Codex's both, to the one listener what is running right now
session files read from local disk: Claude Code's ~/.claude/projects/<slug>/…, OpenCode's session database, Codex's transcripts under $CODEX_HOME or ~/.codex what happened
Claude Code telemetry (optional, from 0.7.1) Claude Code's own OpenTelemetry export, pointed by you at the same loopback listener, and accepted only while agentDeck.telemetry.enabled is on a cost Claude Code estimates, and tool durations

Everything the extension knows comes from those, all local. The live deck is held in memory only and discarded when the window closes. The one thing Agent Deck writes to disk on its own, from 0.7.0, is the stats history: derived numbers, never session content, as append-only JSON Lines under VS Code's own global-storage directory for the extension — never under any engine's directory or your workspace. It is retention-bounded, turned off by agentDeck.stats.enabled, emptied by Clear Stats History, and never read back into a session or a deck.

This matters because "we promise not to send telemetry" is a policy and policies drift. The shipped artifact contains exactly one outbound HTTP call site, and its destination is the loopback literal: a second VS Code window asking the first for its event stream, from 0.7.0 — see §4. Receiving Claude Code's telemetry, from 0.7.1, added no way to send. Zero egress here is a property of what the build contains, and a test fails if that changes.

The extension itself makes no network call. From 0.9.0 the About page offers four links — a portfolio, the source repository, LinkedIn and a sponsor page — and, while no Insights provider is registered, a fifth to the Insights page. Each one asks first ("Agent Deck will open <host> in your browser") and only then is handed to the EDITOR through vscode.env.openExternal, which opens your browser. The extension opens no socket for any of them, fetches nothing, and checks nothing for reachability, so the census in §4 is unchanged by them: one client, one loopback destination. Proof: src/about.test.ts › "every link opens through the editor and nothing else"; src/hooks/egress.test.ts › "contains a server and exactly one client, whose destination is the loopback literal". The urls live in the extension host only: the panel is sent each tile's label and host, and a tile sends back an index. webview/bundle.test.ts › "contains only justified URL literals" holds the panel's bundle to none.

An Insights provider is another extension, and what it hands over is data. From 0.9.0 the extension API (apiVersion 2) lets one extension register as the Insights provider. The provider data is plain JSON, checked field by field, and never executed: Agent Deck reads what the provider returns as own data properties only (never through a getter), refuses a contract member given as a getter, checks every enum against its list, matches every id, version and stats key against a fixed shape, caps every list, and drops and counts what fails.

A provider's text is shown, and every string of it is bounded first. A finding carries the provider's own words: an action (a lead of at most 15 words and a detail), a cause, and a label on each piece of evidence; a refused run carries the step it stopped at and the reason. Each is a name (at most 64 characters), a path (at most 1,024) or free text (at most 2,000) — the first two are the stats history's own caps — and none may carry a control or format character (Unicode Cc and Cf; free text may carry a tab and LF or CRLF line breaks), so no zero-width character and no bidirectional override, nor a line or paragraph separator, nor a lone surrogate. A string that fails is dropped and counted, never shortened. A piece of evidence may be text only on a stats-record field the history itself stores as text, under that field's cap. The panel renders all of it as text, never as markup. Raw output opens as an untitled plain-text document — nothing is written to disk — and more than 1,048,576 characters is not opened at all. Every call Agent Deck makes into a provider is a row of the table below; it asks no licence question and opens no connection for any of them.

Every call Agent Deck makes outside its own code, as of 0.9.0. This is the whole list. A call the extension makes that is not on it is a defect in this document.

call to when
http.request another Agent Deck window's listener, at the loopback literal 127.0.0.1 and agentDeck.port a second window following the first (from 0.7.0). The only network client in the build — §4
vscode.env.openExternal the editor, which opens your browser an About tile (four links) or the Get tile (the Insights page), and only after you press Open on the confirmation that names the host
provider onDidChange(listener) the registered Insights provider once, at registration; the disposable it returns is disposed when the provider leaves
provider listRuns() the registered Insights provider each time the host builds the Insights surface's state, and to check a run is still listed before acting on it
provider getRun(runId) the registered Insights provider for the run you select (the preview), each run you export, and the selected run before raw output is asked for. A set naming any other run is dropped
provider getRawOutput(runId), optional the registered Insights provider only when you press Show raw output on a selected refused run
provider pickAgent(), showPayload(), clearHistory(), optional the registered Insights provider one call when you press that row under Open Insights, and only while the provider has that action
provider investigate(runId), optional (from 0.9.0 round 8) the registered Insights provider one call when you press Investigate Report on a selected run that has a report. It is passed the selected run's id and nothing else; the parent builds no prompt and starts nothing
vscode.commands.executeCommand the editor's workbench.action.openSettings with @ext:nvitlam.agent-deck Open Settings
vscode.commands.executeCommand the editor's workbench.action.evenEditorWidths the deck panel opening while the window has more than one editor group
vscode.commands.executeCommand the editor's setContext, for the three Insights action keys on every change to the registered provider, so the editor's menus offer only the actions it has
vscode.commands.executeCommand Agent Deck's own agentDeck.* commands a sidebar row or a Statistics tab, after the id is checked against the one command table; any other id is dropped
WorkspaceConfiguration.update VS Code's user settings a Tweaks command, for one of Agent Deck's own three settings — §2
window.showSaveDialog, window.showOpenDialog, then workspace.fs.writeFile a file you choose exporting an Insights report — §2
env.clipboard.writeText the clipboard Copy, for one report or for the ticked reports
workspace.openTextDocument an untitled document Show raw output

The stats history (above) is written under the extension's own global storage and is not a call outside it. Two provider members are never called: getLatest (required) and run (optional). Proof: src/hooks/egress.test.ts › "contains a server and exactly one client, whose destination is the loopback literal"; src/extension.test.ts › "a tile ASKS FIRST: the confirmation, then openExternal — DoD 9.23"; src/insights-provider.test.ts › "the provider’s onDidChange reaches the registry’s onChange"; src/insights-provider.test.ts › "DoD 9.44: the selection previews ONLY a run the list holds, read through getRun"; src/insights-provider.test.ts › "DoD 9.44: a getRun answer naming ANOTHER run is dropped, never shown under the row clicked"; src/extension.test.ts › "9.40: the snapshot OFFERS raw output on a placed refused set, and the press opens it untitled and shown"; src/insights-provider.test.ts › "DoD 9.44: invoke() calls the action ONCE, awaits it, and answers what it came to"; src/extension.test.ts › "9.55: the press calls the provider’s investigate with the SELECTED run id and nothing else"; src/insights-provider.test.ts › "a Proxy provider sees no other property touched, and getRawOutput called only when asked"; src/extension.test.ts › "agentDeck.openSettings runs the workbench settings command, filtered to this extension"; src/extension.test.ts › "DoD 4.6c: the deck opens in ViewColumn.One; even-widths runs once with two groups and not with one"; src/extension.test.ts › "9.45: each action is a row under Open Insights ONLY while the provider has it, and the context keys agree"; src/bridge/messages.test.ts › "runCommand accepts every id in the table and NOTHING else"; src/extension.test.ts › "9.47: Copy puts the plain text on the clipboard and opens no dialog".


2. Grounding rules that this posture rests on

These are build-time law in this repository, not guidelines. A change that breaks one fails review.

  • G1 — read-only, always. Amended 2026-09-28: Agent Deck writes only three things: its stats history under globalStorage; export files to a location the user chooses (one report through a save dialog, or one file per ticked report into a folder chosen through a folder dialog); and its own three settings (agentDeck.followNewSessions, agentDeck.openDrawerOnEnter, agentDeck.drawerExpandedByDefault), through VS Code's configuration API to the user settings. Never under any engine's data directory or settings. (This replaces the wording of 2026-09-23, which named the save dialog alone and did not name the settings.) No writes to Claude Code settings, session files, or anything under ~/.claude, ever. Hook installation is a snippet the user pastes themselves; the extension never edits a settings file to install it. In this repository the hook block lives in the repo-local .claude/settings.local.json precisely so that ~/.claude stays untouched.

    Agent Deck's own settings: the Tweaks. From 0.9.0 the sidebar's Tweaks page shows three rows, one per boolean setting — agentDeck.followNewSessions, agentDeck.openDrawerOnEnter and agentDeck.drawerExpandedByDefault. Each row runs one command, agentDeck.tweak.<setting>, and that command writes exactly one thing: the setting's opposite value, through VS Code's WorkspaceConfiguration.update, to the user settings (ConfigurationTarget.Global). It never writes a workspace's .vscode/settings.json, which is usually tracked in your repository. The key is derived from the command's id, so a command that is not a tweak names no key and writes nothing. No webview message carries a key or a value: the sidebar can only ask for a command, and a command id that is not in the one command table is dropped before anything runs. agentDeck.defaultOrdering is not a Tweak; it is a setting you edit in VS Code's settings yourself, and Agent Deck only reads it. Nothing under any engine's directory is written. Proof: src/extension.test.ts › "a tweak command writes through workspace.getConfiguration().update, to Global"; src/view/controls.test.ts › "every tweak command names a declared setting"; src/view/controls.test.ts › "answers NOTHING for a command that is not a tweak"; src/bridge/messages.test.ts › "runCommand accepts every id in the table and NOTHING else".

    A second write, from 0.9.0: an Insights report you export. HTML and Markdown are saved where you choose, through the editor's own save dialog (one file) or folder dialog (one file per ticked report, never replacing a file already there); Copy writes to the clipboard only. The export is built from the checked finding set alone. The HTML page loads nothing — no script, no image, no link, no remote stylesheet, and a Content-Security-Policy of default-src 'none'; style-src 'unsafe-inline' — and the Markdown escapes the provider's text so it cannot become an image or a link. A path inside ~/.claude, the Claude Code projects directory, the Codex directory or OpenCode's data directory is refused, and nothing is written — compared as written and through any junction or symbolic link on the way. A batch into a folder that cannot be listed writes nothing, because "never replacing a file" cannot be kept there. The write lives in one module, src/insights-save.ts, beside that refusal; src/extension.ts still names no write API. Exporting makes no network call. Proof: src/insights-save.test.ts (the one writer and its one importer, the refused roots); src/insights-export.test.ts (every HTML golden parsed for anything that loads; the Markdown escaping; no network import); src/extension.test.ts › "9.47 G1: a path inside an observed engine’s directory is refused and NOTHING is written". src/hooks/listener.ts imports no filesystem API at all, and a test asserts that against the source text — including that it never resolves a home directory.

    Amended 2026-08-27, when a second observation source arrived and measurement contradicted the plain reading. Agent Deck also reads OpenCode's SQLite store, %USERPROFILE%\.local\share\opencode\opencode.db, and that database is in WAL mode. Opening a WAL database read-only writes to SQLite's own -shm shared-memory index, and creates -shm/-wal if they are absent. So G1's claim is stated precisely rather than absolutely:

    No writes to any file the observed engine treats as content.

    opencode.db itself is never modified — measured byte- and mtime-identical across every probe — and auth.json, log/, snapshot/, repos/ and tool-output/ are never opened at all. What is touched is SQLite's lock and index sidecar, which every reader of a WAL database touches, including OpenCode's own process, and which holds no session content. The read-only handle is still the enforcement: a write through it throws ERR_SQLITE_ERROR errcode 8, attempt to write a readonly database.

    The one mode that writes nothing at all, file:…?immutable=1, was rejected for the live database and is used only for this repository's committed test fixtures. It buys zero writes by skipping the WAL, which against a live database means silently returning whatever was last checkpointed — a confidently wrong tree, which is worse than a sidecar. Requesting it on a WAL-mode file is refused in code, so it cannot later be pointed at your data.

    Four secret-bearing tables — account, control_account, credential, session_share — are never read. Five are dropped by schema from any committed fixture: those four plus account_state, which holds no secret itself and exists only to point at account. A superset is the safe direction for a drop list, and the fifth is named here because this document previously implied the two counts were the same one. Dropped by schema means the table is never created in the fixture, so no column named access_token, refresh_token, value or secret exists in the artifact at all — there is nothing to leak even if the drop of a row were ever missed. All five measured zero rows at capture time, which is exactly why the rule keys on the schema rather than on the rows.

  • G2 — source separation. A JSONL parse failure must never take liveness down, and vice versa. The two taps do not share a failure path.

  • G3 — refuse, don't guess. Malformed input increments a counter and is skipped. A schema fingerprint mismatch renders a session unsupported rather than a partial tree. Nothing about input may crash the extension host — see §3.

  • G4 — redaction is production code. Thinking blocks are dropped at the parse boundary, and the signature field is dropped with them: Claude Code writes thinking blocks to disk with an empty thinking string and the bytes in signature, so a redaction that dropped only the visible text would be doing nothing. Tool payloads are truncated with a marker, and large payloads offloaded to tool-results/*.txt go through the same path. Current truncation behaviour, its measured limits and its open items are tracked in the maintainer's working notes; this document does not restate them, because a live description written from inside the phase that is changing them would be wrong by the time it merged.

  • G5 — zero egress. No network except the loopback hook listener. Non-loopback requests are dropped. This is the subject of §4.

  • G6 — fixtures are law. Parser behaviour is pinned to bytes captured from real sessions.

The second read-only source, stated in full

G1 above says what is never touched. This says what is, because "read-only" is a claim about scope as much as about direction, and a reader cannot check a scope that is only ever described by its complement.

Where. %USERPROFILE%\.local\share\opencode\opencode.db — one SQLite file, opened through node:sqlite's DatabaseSync with { readOnly: true }. AGENT_DECK_OPENCODE_ROOT overrides the directory, which is how every test reaches a fixture instead of your data. Nothing else in that directory is opened: not auth.json, not log/, snapshot/, repos/ or tool-output/.

What. Six tables, and only these columns. The engine asserts every one of them before it reads anything; a missing table or column refuses the store outright rather than rendering a partial tree (G3). Unknown tables and columns are ignored and counted, never read.

Table Columns read
project id, worktree, vcs
session id, project_id, parent_id, version, agent, title, directory, slug, model, cost, tokens_input, tokens_output, tokens_reasoning, tokens_cache_read, tokens_cache_write, time_created, time_updated, time_archived
message id, session_id, time_created, time_updated, data
part id, message_id, session_id, time_created, time_updated, data
event id, aggregate_id, seq, type, data
event_sequence aggregate_id, seq, owner_id

Three of those session columns — tokens_reasoning, tokens_cache_read, tokens_cache_write — are asserted but never selected. The schema contract names them, so if a future OpenCode drops one the engine refuses instead of quietly rendering a tree built on a shape it has never seen. A required column that nothing reads looks like an oversight and is not.

Reasoning content is dropped at the parse boundary, before any record is built — the G4 rule, applied to the second engine. For OpenCode the reasoning bytes exist verbatim in the store, so unlike the Claude Code case that test cannot be vacuous: it searches the produced SessionState for the literal captured bytes.

No SQL from a caller, ever. Every statement is a fixed SELECT written in src/opencode/db.ts; no query is assembled from input, and no database handle escapes that module.

How it degrades, without ever crashing (G3). A missing file, an unreadable one, or a corrupt one each surface as a named degrade — databaseMissing, databaseUnreadable, databaseCorrupt — and leave Claude Code sessions rendering unchanged (G2). A schema that is not OpenCode's renders every session unsupported. A graft that cannot place a row surfaces as graftFailed, which is a containment rather than a fix: it keeps a throw from escaping into the extension host, at the cost of darkening every OpenCode session over one unplaceable row. That trade is recorded rather than presented as a solution.

The third read-only source, stated in full

Same treatment as OpenCode above, and for the same reason: a scope described only by its complement is a scope nobody can check.

Where. $CODEX_HOME when it is set and non-empty, otherwise ~/.codex — resolved at read time, never captured at module load. That variable relocates Codex's entire surface, credentials included, so an engine that hard-coded the home location would observe nothing at all for such a user while reporting a confident absence. An explicit root passed by the caller outranks both, and is how every test reaches a fixture instead of your data (G6).

What, exactly. Four operations and no others:

Path How it is read
<root>/sessions/**/rollout-*.jsonl byte-offset tailing through the shared FileTail
<root>/sessions/** readdirSync to discover those files
<root>/thread-writer-locks/ readdirSync — names only
a transcript statSync().mtimeMs, for the liveness fallback

Time and spawn results (0.8.0) add no content to a stats record. F14's figures are differences and ratios of the timestamps the engine wrote on each tool call and usage turn — never a clock reading and never message text. F15 reads one structural fact, the status of the call that spawned an agent, and no preview. A session read in part is excluded from stats entirely, as a parked one is. Proof: src/stats/timing.test.ts › "states a wall time that is the span of its own instants, not its envelope"; src/stats/f15.test.ts › "reads the status and no preview — an errored spawn has its result".

A transcript over agentDeck.codex.maxTranscriptBytes (0.8.0) is read in two windows and no others: its first 256 KiB, on which the fingerprint and the workspace match run, and its last 16 MiB. The bytes between are not read. The session is marked partial on the deck and on SessionState.partial, and the line the jump lands in is dropped and counted, never parsed. Proof: src/perf/oversize.test.ts › "reads 16.25 MiB of it, in those two windows and nowhere else"; src/codex/index.test.ts › "reads 65 MiB as 256 KiB of head plus the last 16 MiB, and says so".

The lock files are never opened. They are 0 bytes and their whole content is their name, so the engine reads the directory listing and stops there. .coordination.lock is process-lifetime and is not a thread; it is excluded by name.

Neither hooks.json nor config.toml is opened by this extension. They are Codex's own files; you edit one and Codex's trust prompt covers the other. Agent Deck never reads them, and G1 already forbids writing them.

Discovery walks the tree; it never composes a path from a clock. Rollout files are partitioned by the day a thread started, so a session running past midnight puts a child under a different day from its parent, and a reader that built YYYY/MM/DD from Date.now() would silently miss it.

No socket to Codex. No App Server, no app-server proxy, no second port. Codex hooks POST to the same loopback listener Claude Code's do — one socket for the whole extension, which is what §4 audits.

Reasoning is dropped at the parse boundary (G4), and for Codex there are two shapes of it: response_item records typed reasoning including encrypted_content, and event_msg items typed Reasoning including summary_text and raw_content. Both the plaintext summary and the encrypted bytes are dropped, never stored, never decoded, never displayed. A spawned agent's task description is encrypted in the hook payload and is likewise never decoded — its task name arrives in plaintext and is what labels the node, which is a real and deliberate asymmetry rather than an oversight.

Large tool output is stored WHOLE and INLINE by Codex — no offload file, unlike Claude Code. 248,000 bytes of stdout have been measured in a single record. The hazard here is therefore a very long line rather than missing content, and truncation is applied on our side before anything is displayed.

G10 — the never-opened list

Named in code as an exclusion list, not filtered after the fact: the name is judged before any path is joined, stat'ed or descended into, so there is no moment at which one of these exists as a string that has been handed to the filesystem. src/codex/never-open.ts holds the list, and a test greps the engine's source for each name and asserts it appears only there.

Never opened
files auth.json, installation_id, cap_sid, models_cache.json
directories .sandbox-secrets/** — not descended into, not stat'ed inside, not reported
suffixes *.sqlite, *.sqlite-wal, *.sqlite-shm

The SQLite exclusion is by decision, not by difficulty. Those are live databases Codex is writing. The read-only-WAL argument that lets the OpenCode engine open its store does not transfer for free, and opening one is a separate question with its own gate.

The trust step is yours, and it is manual

Codex requires a hook command to be trusted before it will run, and that click happens in the Codex extension — not here. Two consequences worth knowing:

  • Editing hooks.json invalidates the trust entry, and the hook then silently stops firing until a human re-trusts it. If liveness goes quiet after you change that file, this is why.
  • Six events sharing one identical command produce six distinct trust hashes, one per event. That is measured; the mechanism is not, so do not infer one. A check written expecting a single hash across six events reports a failure that is not there.

3. The listener's trust boundary

src/hooks/listener.ts is the only inbound surface. Its properties, each followed by the test that proves it — a file and the exact title of a test in it, which src/release/surfaces.test.ts checks exist:

  • Binds the literal 127.0.0.1, hard-coded, never a hostname and never the wildcard. The bind host is a module constant and is not configurable; only the port is. Proof: src/hooks/listener.test.ts › "binds to the literal loopback address, never a wildcard"; src/hooks/egress.test.ts › "binds the loopback literal and never a wildcard, in the built artifact".
  • Validates the socket's remote address on every request. Proxy headers — X-Forwarded-For, X-Real-IP, Forwarded — are attacker-controlled strings and are never consulted for that decision. A non-loopback origin is answered 403 and counted. Proof: src/hooks/listener.test.ts › "drops a non-loopback POST with 403 and counts it"; src/hooks/listener.test.ts › "does not let proxy headers grant loopback status".
  • Refuses an ephemeral port. The pasted hook snippet names a fixed port literally, and there is no discovery file to tell it otherwise — writing one would break G1. A port collision surfaces as a typed error the user is shown; the listener never silently rebinds somewhere else. The one way to bind port 0 is an option marked TEST-ONLY in the source, which exists so the suite can bind a port atomically instead of racing for one; a source scan asserts that no production module under src/ names it, and it cannot change the bind address either way. Proof: src/hooks/listener.test.ts › "refuses an ephemeral port rather than binding one"; src/hooks/listener.test.ts › "surfaces a port collision as a typed error and does not rebind elsewhere"; src/hooks/listener.test.ts › "no production module outside listener.ts names either TEST-ONLY option".
  • Caps request bodies, by two guards, because one of them cannot see the other's cases. The cap is DEFAULT_MAX_BODY_BYTES in src/hooks/listener.ts. A body whose declared Content-Length already exceeds the cap is never buffered at all — that is an allocation guard on the headers alone. A body that arrives with no declared length, i.e. Transfer-Encoding: chunked, is measured as it streams and cut at the same cap. Either way the request is counted and answered 413. A body that keeps arriving past a hard multiple of the cap has its socket destroyed rather than held open. Proof: src/hooks/listener.test.ts › "a body over the size cap → 413"; src/hooks/listener.test.ts › "a chunked body with no Content-Length is capped WHILE it streams".
  • Never throws on input. Every refusal path increments a named counter and answers a status code. Consumer callbacks that throw are caught and counted, so a downstream bug cannot take the socket down. Proof: src/hooks/listener.test.ts › "survives the whole malformed battery back to back"; src/hooks/listener.test.ts › "a consumer that throws is counted and does not break the response".
  • Serves no files and reads no paths. It answers exactly six paths: the hook event path; /agent-deck/identity and /agent-deck/events, which only another Agent Deck window on this machine asks for (from 0.7.0); and /v1/metrics, /v1/logs and /v1/traces, Claude Code's telemetry export (from 0.7.1). Every other path is a 404, including traversal attempts, which have nothing to traverse to. Proof: src/hooks/listener.test.ts › "G1: the listener imports no filesystem API and writes nothing"; src/hooks/listener.test.ts › "an unknown path → 404"; src/hooks/telemetry-route.test.ts › "a neighbouring path is a plain 404 on the event path accounting".
  • Refuses Claude Code telemetry unless agentDeck.telemetry.enabled is on (default off): 403, and the body is drained and never parsed. When it is on, a body must say it is JSON (415 otherwise), be an OTLP body for that path (400) and fit the same cap (413). It is parsed ONCE, where it arrives, by an allow-list: the five account attributes Claude Code attaches to every record and the prompt and response fields are dropped there, so they never reach the session model, the relay to other windows, the stats history or the diagnostics channel. No answer is retryable. Proof: src/hooks/telemetry-route.test.ts › "parses NO body while off: malformed and oversize bodies are 403, not 400 or 413"; src/otel/parse.test.ts › "lets no attribute NAME and no placeholder VALUE survive into the slice"; src/otel/parse.test.ts › "drops prompt, response and user_prompt when they hold actual prose"; src/hooks/shared.test.ts › "no OTLP body and no identity attribute crosses the wire".

The three telemetry routes, and every answer they give. /v1/metrics, /v1/logs and /v1/traces, OTLP over HTTP in JSON, on 127.0.0.1 at agentDeck.port. A non-loopback origin is refused before any route is chosen; the method is checked first and the content type before the setting, so a protobuf body sent while the setting is off is a 415, not a 403. Proof: src/hooks/telemetry-route.test.ts › "a non-loopback origin on a telemetry path is the plain 403 drop, counted as such and nowhere else"; src/hooks/telemetry-route.test.ts › "checks the content type BEFORE the setting: protobuf while off is 415, not 403".

Answer When Proof
200 Accepted: a POST, a JSON content type, the setting on, an OTLP JSON body for that path, within the cap. Parsed once and published. src/hooks/telemetry-route.test.ts › "DoD 6.2 — 200: every captured body is accepted, parsed once and published"
400 JSON, but not an OTLP body for that path: empty, malformed, not an object, or another signal's body. src/hooks/telemetry-route.test.ts › "DoD 6.2 — 400: JSON, but not an OTLP body for this path"
403 agentDeck.telemetry.enabled is off. The body is drained, never parsed, and the answer names the setting. src/hooks/telemetry-route.test.ts › "DoD 6.2 — 403: the setting is off (the shipped default)"
405 Any method but POST, whatever the setting. src/hooks/telemetry-route.test.ts › "DoD 6.2 — 405: any method but POST, whatever the setting"
413 Over DEFAULT_MAX_BODY_BYTES, 512 KiB — the hooks' cap, declared or streamed. src/hooks/telemetry-route.test.ts › "DoD 6.2 — 413: the hooks cap, unchanged, at limit and limit+1"
415 A content type that does not say JSON, or none at all. src/hooks/telemetry-route.test.ts › "DoD 6.2 — 415: a POST that does not say it is JSON"

No answer asks the exporter to retry: none of the six is one of the statuses OTLP over HTTP retries on (429, 502, 503, 504). Proof: src/release/surfaces.test.ts › "no answer in the table is a status OTLP retries on".

A 403, 405 or 415 is sent after the last byte of the body has arrived, so a large body reads the refusal rather than a connection reset. Proof: src/hooks/telemetry-route.test.ts › "DoD 6.2 — a refusal drains the body first, so the refusal is what arrives".

Hostile-input testing. fixtures/synthetic-hook-fuzz/corpus.jsonl is a synthetic corpus replayed over a real loopback socket against a real listener at the shipped default body cap. It covers malformed JSON, bodies truncated mid-token, oversized bodies and the exact byte at the cap boundary, wrong and absent content types, wrong methods and routes, raw C0 control bytes versus the same bytes as \u escapes, invalid UTF-8, BOM prefixes, lone surrogates, __proto__ and constructor.prototype bodies, deeply nested JSON, type-confused fields, unknown forward-compatible fields, minimal-but-valid payloads, and spoofed off-box origins. Each case asserts the status code and the exact counter deltas — every counter not named must be unchanged — because "it did not crash" is satisfied by a listener that answers 200 to everything.

The corpus speaks through an HTTP client, so every one of its bodies arrives with a truthful Content-Length and every oversize case is therefore answered by the declared-size guard. Reaching the streaming guard needs a request the corpus format cannot express, so those cases are driven from a bare socket in src/hooks/listener.test.ts: a Transfer-Encoding: chunked body with no declared length at all, an understated Content-Length (measured: node frames the body by the declared length, so the surplus becomes the next request rather than reaching the cap), an overstated Content-Length, an unparseable request line, a header block that never terminates, and pipelined requests. See that directory's README.md.

What the boundary does NOT protect against, stated plainly: any process running as you on your own machine can POST to the loopback port and inject fabricated liveness events. There is no authentication, and adding one would mean a shared secret that the pasted hook snippet would have to carry. The consequence of that injection is bounded by what the extension does with an event: it renders it. Nothing is executed, nothing is written, nothing leaves the machine. An attacker already running code as you has far better options than lying to a read-only panel.


4. The zero-egress audit

Three parts, all in src/hooks/egress.test.ts, all run in the ordinary suite. None can skip. §4a and §4b below describe the first two, which cover the shipped host bundle.

The third covers the OpenCode engine, and it exists because the first two cannot: it bundles src/opencode/index.ts as its own entry point and denies node:http as well as dns and net. The host bundle cannot make that claim, because there the loopback hook listener is the one sanctioned socket — so a host-bundle scan would pass while an engine that opened an HTTP client hid behind the listener's allowance. The OpenCode engine opens zero sockets of any kind. Its liveness is a cursor over the event_sequence table, not a subscription; the SSE accelerator OpenCode offers is deliberately not used. The same describe asserts the shipped bundle never reaches the test-only fixture builder, which is the one module in that tree that opens a database for writing. Proof: src/hooks/egress.test.ts › "reaches no network-capable module at all — not even node:http"; src/hooks/egress.test.ts › "never reaches synthetic.ts, the one module here that can write".

4a. Dependency review — what could open a socket

Method. The VSIX ships dist/ and nothing else: vsce is invoked with --no-dependencies and .vscodeignore excludes node_modules/**. So the shipped runtime surface is the esbuild bundle, not a dependency tree that never gets installed on a user's machine. The audit therefore enumerates the module ids the built bundle actually requires and gates them, rather than auditing a lockfile. Proof: src/hooks/egress.test.ts › "the shipped artifact is the bundle, not a node_modules tree".

What is asserted, on a bundle built on demand rather than read off disk (the test builds its own rather than trusting whatever dist/ holds when it runs, which could silently be an old one):

  • every module id the bundle names is either a node: builtin or vscode — nothing third-party survives bundling as a separate module. Both spellings of "names" are scanned: require("x") and dynamic import("x"), which esbuild leaves verbatim rather than rewriting, so a require-only scan would have been blind to it; Proof: src/hooks/egress.test.ts › "names only node builtins and vscode — no third-party module survives bundling".
  • none of net, tls, https, http2, dns, dgram, child_process, worker_threads, cluster or inspector is reachable, in either the bare or node: spelling, through either form. That is proved by injection rather than asserted: the same check applied to the real bundle with one dynamic import appended must report the injected module; Proof: src/hooks/egress.test.ts › "reaches no network-capable module other than the listener"; src/hooks/egress.test.ts › "the scan sees a dynamic import(), proven by injecting one".
  • node:http is present, so the check is not vacuous — it is the listener; Proof: src/hooks/egress.test.ts › "reaches no network-capable module other than the listener".
  • exactly one outbound call site is compiled in — the follower's http.request, whose host is the loopback constant the server binds (from 0.7.0) — and no other client API: no https.request, no fetch(, no XMLHttpRequest, no new WebSocket(, no navigator.sendBeacon. From 0.7.1 the same bundle carries the telemetry parse and join, and the only address literal in it is 127.0.0.1; Proof: src/hooks/egress.test.ts › "contains a server and exactly one client, whose destination is the loopback literal"; src/hooks/egress.test.ts › "carries the telemetry route and join, and no new way to send (DoD 6.5)".
  • the loopback literal appears and 0.0.0.0 does not, asserted against the built artifact so that a build step rewriting a constant could not slip past the source-level guard. Proof: src/hooks/egress.test.ts › "binds the loopback literal and never a wildcard, in the built artifact".

Limits, honestly. Nobody read every dependency's source. This is a reachability argument about the shipped bundle plus the runtime measurement below — not a proof that no dependency contains egress code somewhere. The webview is a separate artifact with its own bundle guard and its own strict CSP, and is not covered by this section.

4b. Runtime socket census — what actually opens

Method. A child node process stages the freshly built dist/extension.cjs beside a stub vscode module (module resolution is relative, so the stub is what require('vscode') finds — the same technique the extension-host load test uses). Before the bundle is loaded, the child wraps net.Socket.prototype.connect, net.connect, http.request/get, https.request/get, tls.connect, dgram.createSocket, the dns resolvers and Module._load. It then drives the real exported activate() against the committed fixtures, and counts live handles with process._getActiveHandles() at four points, classifying each by its underlying libuv handle — TCP is the only kind that can leave the machine.

The harness's own probe POST deliberately speaks raw HTTP over a socket connected with the pre-instrumentation connect, so it does not appear in its own census; that it returned 200 proves the connection really happened and the listener really answered.

Result (Node v24.15.0, Windows, one run; reproduce with AGENT_DECK_CENSUS_DEBUG=1 npx vitest run src/hooks/egress.test.ts):

phase handles observed
bundle loaded, before activate() none at all
after activate() Server/TCP on 127.0.0.1:<configured port>, plus 19 FSWatcher/FSEvent
after serving one hook event the same, plus the harness's client socket and the child's stdout pipe
after deactivate() no TCP handle, no watcher

Outbound connections attempted by the extension: 0 at load, 0 through activation, 0 across the entire run. The file watchers are FSEvent handles — local filesystem, not sockets. Proof: src/hooks/egress.test.ts › "has no socket before activation and none after disposal"; src/hooks/egress.test.ts › "attempts zero outbound connections from load through activation"; src/hooks/egress.test.ts › "attempts zero outbound connections across the entire run"; src/hooks/egress.test.ts › "opens exactly one socket with three engines live: the loopback listener".

One measured surprise, recorded rather than smoothed over. The run is not DNS-silent. Node's own Server.listen(port, host) routes through lookupAndListen → dns.lookup(host, { all: true }) even when the host is already a literal IP address. So there is exactly one DNS call in the whole run, its argument is the string 127.0.0.1, and its caller — captured from the stack — is the inbound bind. The test asserts all three: the count is exactly one, the argument is a loopback literal, and the stack names the bind. Asserting the run is DNS-free would be false, and asserting merely "at least one lookup, and each looked fine" would be weaker than this paragraph claims — which is how a measured finding turns into a comfortable story. Proof: src/hooks/egress.test.ts › "resolves no hostname: every DNS call is node resolving the loopback literal it was told to BIND".

Limits, honestly. This measures one activation cycle on committed fixtures, on one OS and one Node version. Code paths that cycle never exercises are not covered by it. It measures the Node extension host, not the webview.


5. Installing the hook

Hook installation is a manual paste block; the extension never writes it for you (G1). Two things about the command in that block matter for your own safety rather than ours:

  • It must fail fast when nothing is listening, because it runs inside your real session — now your Codex sessions as well as your Claude Code ones, since both engines post to this one listener. The block uses node -e rather than curl: node takes ECONNREFUSED and exits 0.

    Re-measured 2026-09-04 against a closed loopback port, five runs each, timing the exact one-liner this README pastes — read out of the README rather than retyped, so the number describes what ships:

    command min median max exit
    the shipped node -e block 87 ms 89 ms 99 ms 0
    curl.exe 8.18.0, -m 5 2,154 ms 2,158 ms 2,169 ms 7

    A ratio of about 24×, and the number that matters is the median, not the ratio: 89 ms is a cost you would not notice on a tool call and 2.2 s is one you would, on every tool call, in a session you are trying to work in.

    Two earlier measurements on this same machine are kept rather than overwritten, because the spread is the point: node -e at 81 ms / exit 0, and curl.exe at both 2,098 ms / exit 7 and ~1,140 ms / exit 28. The exit code differs with whether the port is refused or filtered and the timing differs with the curl build; no measurement has ever put them within an order of magnitude of each other, which is the claim the block rests on.

  • The POST is unconditional. With nothing bound, it is refused and nothing happens. Do not read a quiet listener as evidence that hooks have stopped firing.


6. Hard exclusions

Not implemented, and not accepted as contributions: writes to anything an observed engine owns · replay of a session, or persistence of its content · wrapping or launching any observed engine · sending telemetry, or any egress. The stats history in §1 is the one write Agent Deck makes on its own, and it holds derived numbers only, in the extension's own storage; an Insights export (§2, G1) is written only where you choose, and never into a directory an observed engine owns; a Tweak (§2, G1) writes the one Agent Deck setting you toggled, to your user settings. Zero writes to what is observed is the trust anchor, and the point of writing it down is that it is easier to defend a boundary than to relocate one.

There aren't any published security advisories