Skip to content

Commit 0060242

Browse files
committed
Merge master (#374, #362, #369) into feat/p-legible-1-defender-manifest
desktop/main.ts: the engine env keeps master's meetingsEnv and this PR's LUCID_HOST_EXE. PROGRESS.md: master's entries, then this PR's P-LEGIBLE.1 entry.
2 parents cf32540 + 97acae1 commit 0060242

47 files changed

Lines changed: 8036 additions & 129 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎DECISIONS.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -24147,6 +24147,14 @@ P-SANDBOX.13 designed around this (the Add folder route never accepts a path fro
2414724147
- The broker serves the workspace only, not user-granted folders.
2414824148
- Not yet exercised: a real host run from a contained session. The agent session that built this is itself inside the AppContainer, so its test engine hit the same getcwd wall at the validation step. The shim round trip and the live route's refusals were verified.
2414924149

24150+
## ADR-0385 -- P-RECOVER.1: LUCID recovers itself, and says so with an incident report (2026-09-23)
24151+
24152+
**Context.** Beta users reported three symptoms: the IDE fails to connect to an agent or loses it mid-session; a prompt shows "reconnecting" and nothing happens; closing and reopening the IDE does not help until the user kills the LucidAgentIDE or omp process. The field logs (engine.log, ~/.omp/logs/omp.*.log, lucid-acp.log) and the source gave five causes. (1) `Backend.prompt()` refused with `A chat turn is already running` 22 times while `/api/chat/status` said idle: `loadSession()`/`newSession()` cleared the turn record but left its listener installed and its `session/prompt` pending, so every later prompt was refused until the health ladder respawned the child about 7 minutes later (reproduced against the real Backend and a hanging fake agent). (2) The MASTER omp connection had no death handling: `start()` returned early on a dead `this.acp`, so after an omp crash every prompt failed with "agent process exited" until the 30 s health tick, and after two recoveries in one episode the ladder pinned at "needs a manual restart" with nothing on screen. (3) Closing the main window did not quit the app while the agent browser window was open (`window-all-closed` needs every window closed), and `second-instance` only focused an existing `win`, so a relaunch silently did nothing while the headless app held the single-instance lock. (4) An engine orphaned by a crashed main squatted 5319 (ADR-0381/0382 cover the engine; nothing covered its children or the omp grandchild, which on Windows is a `bun.exe cli.js` behind the `node_modules/.bin/omp.exe` shim that `proc.kill()` does not reach). (5) omp died with an uncaught `EPIPE: broken pipe, write` from the security gate's stderr notice after the engine holding the read end exited (omp fails closed on a handler throw, so no bypass, but the agent process was lost).
24153+
24154+
**Decision.** Recover in place wherever ownership is provable, and leave a redacted incident report every time. Startup: a run ledger (`<userData>/run-state.json`) marks a clean exit; after an unclean one, main enumerates processes once and stops ONLY what the ledger proves is ours (the recorded engine pid with the same image path and a start time within 5 s, its descendants by creation-ordered parent edges, and orphans whose parent is the dead recorded engine), never by name, never the current main; one `taskkill /F` names every proven pid (no `/T`, whose parent walk has no creation-order check). The engine persists the master session id (`~/.omp/lucid-last-session-<PORT>.json`), snapshots the previous one at start, and the window resumes it with a VERIFIED `session/load` (failure is reported, a fresh session starts, and the user is told "Your previous session could not be recovered. A new session was started."). Runtime: a cleared turn releases its own listener and cancels its session; a dead master child is revived on demand with the same session id; `ACPClient.stop()` tree-kills on Windows so tool processes the agent started do not outlive it, and resolves true only once that is confirmed (taskkill exits 0 and the child's own exit is observed; POSIX: the child exits after SIGTERM, then SIGKILL); a replacement master is never spawned while the child `restart()` retired is unconfirmed, so a watchdog or window recovery that cannot confirm it fails with nothing spawned and nothing re-sent; a deliberate exit (quit, the settings relaunch, the GPU relaunch, the foreign-port quit) marks the run ledger clean only after the engine and everything under it, the omp tree included, are verified stopped (bounded at 20 s), and otherwise leaves it unclean for the next launch's reaper; the renderer runs a bounded supervisor on "reconnecting" or a refused send (probe, reattach, recover the agent once, restart the engine once through a main-process IPC that refuses while the engine's nonce health still answers and is rate-limited, then give up with a plain message). Incidents live in `<userData>/incidents/<id>.md|.json` (main writes there directly and hands the engine the same folder as `LUCID_DATA_ROOT`), outside the `~/.omp` tree the contained agent may write; with no data root outside it, nothing is recorded. Every text is redacted on the way in (the support-bundle rules plus e-mail, long hex, and profile paths), the home path itself is never stored, developer `[ASKSAGE_DIAG]` records (whose `raw` field quotes provider output) are stripped from every log tail, which is read from a line boundary, and incidents never contain prompts, transcripts, settings or credential stores. Every read validates the stored record field by field and rebuilds the title, the public issue body and the report from it, so nothing shown or prefilled is read back from disk as text. Submitting is always the user's action: the prefilled GitHub issue carries the summary only, because the repository is public; the full report stays local for the user to review and attach.
24155+
24156+
**Consequences.** The logged wedge and dead-child paths self-heal in seconds instead of minutes or never, a relaunch always yields a window, and every recovery leaves evidence a maintainer can act on. Automatic stopping of leftovers is new policy: it is limited to ledger-proven processes, and anything the ledger cannot prove still goes through the ADR-0382 warning dialog. If a retired agent's tree cannot be confirmed stopped, that engine refuses to start another agent (each attempt retries the stop while the old root still runs) and says to quit and reopen LUCID: availability is traded for never running two agents in one workspace. Quitting waits for the engine tree to be stopped and verified. A standalone engine without `LUCID_DATA_ROOT` records no incidents. Open for maintainers: whether the public issue tracker is the right destination for incident summaries or a private channel (support@) should be offered instead; and the Agent Builder / scheduled agent path still runs `omp -p` through `Bun.spawnSync`, which blocks the engine event loop for up to 120 s per segment and, on timeout, kills only the shim (the real omp keeps working): a separate increment.
24157+
2415024158
## ADR-0384 -- P-LEGIBLE.1: legible to Defender and Agent 365 without a content path (2026-09-23, issue #302)
2415124159

2415224160
**Context.** Microsoft Defender for Endpoint now inventories local AI agents (portal: Assets > AI agents > Local agents; advanced hunting: `AgentsInfo | where Platform == "LocalAgents"`) and, in preview, protects them at runtime. In an M365 E7 / Agent 365 shop an agent that is not in that inventory reads as shadow AI and risks an endpoint block regardless of merit. The issue asked what "legible" concretely requires, whether it fits our invariants, and for the minimal increment. Findings from Microsoft's primary docs (discover-local-ai-agents, local-agent-discovery-overview, ai-agent-runtime-protection-overview, all dated 2026-09-16): (1) Discovery is a Microsoft-maintained list of supported agents (Claude Code, Codex CLI, Cursor, Copilot and roughly 30 more); there is no public manifest or registration API a third party can use to enroll. The profile Defender builds is `Name`, `Version`, `McpServers`, `DeclaredTools` and `RawAgentInfo.localAgentMetadata` = `vendor`, `relatedProcess`, `trustedProcess`, `autoApprove` (both reported as the strings "true"/"false"), device, account, `localMcps`. (2) Entra Agent ID does not apply: local agents resolve to the OS user through a `used by` edge; `can authenticate as` is for cloud agents. (3) Runtime protection hooks three checkpoints (user prompt, pre-tool call, post-tool response) through each vendor's own hook interface (Claude Code, Codex CLI, Copilot CLI hooks). The fallback, network inspection, does not support certificate pinning or HTTP/3 and by construction cannot see loopback traffic, so for this IDE it observes essentially nothing. (4) In audit or block mode, detections forward to Defender XDR in the cloud, and Defender discovery requires the commercial cloud (sovereign and national clouds are not supported), which matters for CUI deployments. (5) `trustedProcess` describes the host binary; `build-desktop.yml` documents our installers as unsigned (it signs only when the `WIN_CSC_LINK` / `MAC_CSC_LINK` secrets exist, and macOS otherwise builds with `identity: null`), so a future profile would likely report "false".

‎PROGRESS.md‎

Lines changed: 11 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,6 +4,13 @@ Three lines per session: **shipped / stubbed / next** (CLAUDE.md session ritual)
44

55
-----
66

7+
## P-RECOVER.1: self-recovery at startup and mid-session, with incident reports (ADR-0385)
8+
- **shipped:** session-switch wedge fix (the logged "A chat turn is already running"), on-demand revival of a dead master omp with the same session, Windows tree-kill in ACPClient.stop(), EPIPE-proof gate notices; run ledger + ledger-proven leftover reaper at launch (live-proven on Windows: only the recorded tree died, a same-image stranger survived), verified resume of the previous session, second-instance always yields a window, closing the main window quits; bounded renderer supervisor (probe, reattach, recover agent, restart engine via guarded IPC, give up) and startup/incident notices with a summary-only public-issue submit dialog; incident store with redaction. Tests: incident 7, engine recovery 15, ledger+reaper 20, supervisor 26; wedge, tree-kill and EPIPE tests each fail on the pre-fix code. Full suite 5606/0, three typechecks clean. Smoke: an isolated source engine on 5391 resumed the seeded previous session, settled the startup incident, revived a killed agent in place, and the renderer showed the notice, the submit dialog, and a terminating give-up when the engine died.
9+
- **review fixes (merged with master e51df61):** incidents moved out of the agent-writable `~/.omp` to `<userData>/incidents` (`LUCID_DATA_ROOT` for the engine; none recorded without it), and every read validates the record and rebuilds the title, public issue body and report from it; the home path is never stored; `[ASKSAGE_DIAG]` records and `"raw":` fragments are stripped from every log tail, read from a line boundary; `ACPClient.stop()` now resolves true only once the tree is confirmed ended (taskkill exit 0 plus the child's exit), and the backend never spawns a replacement while the retired child is unconfirmed, so a recovery that cannot confirm it fails with nothing re-sent; every deliberate exit marks the run ledger clean only after the engine tree (omp included) is verified stopped.
10+
- **stubbed:** nothing. Not exercised in a real Electron run: second-instance reopen, close-with-agent-window quit, the engineRestart IPC end to end, and the verified-stop quit path (unit-level and smoke coverage only).
11+
- **next:** decide the incident destination (public issues vs private channel); move agent_run.ts off Bun.spawnSync and tree-kill its timeout.
12+
13+
714
## P-MASCOT.4: themed roster + expressive eyes + Regular/Max tiers + the run-only arcade
815
- **shipped:** six themed 40x52 palettes (lucid/ember/glacier/orchid/solar/stealth) with original expressive eyes and three new work activities (coding, scanning, staff) rotating costumes in sync; the Agent role's Regular/Max model tier (provider-route + family preserving, catalog-grounded via model_families, CUI/managed lockdown and Fable credential gates honored, confirmed switches with rollback + send-hold, exit restores the prior model only after backend confirmation); the opt-in under-composer arcade easter egg (deterministic pure engine, crates/beams/pots/targets + 1.7x ninja stars, debris, Flycatch chopstick bonus every 5 clears, level speed curve capped 250) whose points join the persistent Trivia Wire balance as shared LUCID points (awardBonus on the same lucid.trivia tally); demo-P-MASCOT.4 runs 110 tests + the avatar-4 demo; ADR-0260 served-byte gate green (all new markers present, retired resolveConversationModel absent).
916
- **stubbed:** the host app's Preview panel wedged on a stale boot-time bundle mid-QA, so the ENRICHED arcade art (backdrop/pots/stars/flycatch scene) is verified by engine tests + painter review + auto-demo harness, not a fresh screenshot; the QA auto-demo (.mascot-qa/arcade.ts) drives real buttons/keys on a timer and stays in the repo for the next visual pass.
@@ -5073,7 +5080,10 @@ Roadmap phases (each its own future increment + ADR for its frozen-contract delt
50735080
- **shipped:** `tools/git-broker/git.cmd` + `git_shim.ts` (first on the contained PATH) POST to `/api/git/exec` (agent token); `desktop/git_broker.ts` allowlists subcommands and options, confines path arguments to the workspace (junction-aware), validates `.git/config` against a key allowlist while holding it and `.git` open without write/delete sharing, forces no-hooks/no-fsmonitor/https-only overrides, and routes network git through the egress proxy; `push -u` and tracking checkouts set upstream after the held window. `make demo-P-SANDBOX.17` (share modes proven on NTFS, shim round trip); a booted engine's route returns 403 without a token and legible refusals.
50745081
- **stubbed:** no host run from a contained session yet (this session's own engine is contained, so git stops at getcwd during validation); ssh remotes, stdin, worktrees and submodules unsupported by design.
50755082
- **next:** rebuild and install, then run `git status` / commit / push from a contained session; then the three small PRs (#379, #372, #358).
5076-
5083+
## P-MEET.1: the Meetings panel, a thin client of the Lucid Meeting Hub (LUCIDMeetingHub#7)
5084+
- **shipped:** finding "what did we decide last Tuesday" meant leaving the IDE, because meeting recall existed only conversationally through `lucid_tool.py`. New left-rail Meetings fly-out: search, recent list (title/date/app), detail (summary, decisions, open action items), and an upcoming-meeting row that deep-links the Hub window. `desktop/meetings_hub.ts` is the engine-side client of the Hub's bearer-scoped `/ext/*` surface on loopback, 127.0.0.1:5123 by default (`/ext/meetings`, `/ext/meeting/<file>`, `/ext/todos?open=1`, `POST /ext/todos/mark`, `/ext/premeeting`, `/ext/pair/claim`) with every field re-typed at the boundary; `desktop/renderer/meetings_panel.ts` is the pure HTML half. The pairing bearer goes to the OS-encrypted vault under ref `meeting_hub_token` (`main.ts` injects it as `LUCID_MEETING_HUB_TOKEN`, the Figma PAT seam) and never to the settings file. Dormant Hub = one honest info row behind a 300ms HEAD probe, no retry loop, no polling anywhere. A locked Hub vault renders metadata rows plus a notice, never "no meetings". Review fixes: the engine captures `LUCID_MEETING_HUB_TOKEN` into module state and deletes it from process.env when meetings_hub.ts loads (before any omp/fleet/scanner child spawns, the P-SANDBOX.15 discipline); `LUCID_MEETING_HUB_URL` is refused unless it is plain `http://` on 127.0.0.1, localhost or [::1] (any port), and a refused origin is contacted by nothing; a failed `/ext/todos` read leaves action items UNKNOWN (a "?" line plus a notice), never struck through as done; both Hub links use the engine-validated origin, so a configured port is honoured; the pairing and success copy says the token can read meetings AND mark action items done. 55 tests across the two new suites (`bun test desktop/meetings_hub.test.ts desktop/renderer/meetings_panel.test.ts`), including a fresh-process check that the engine still authenticates while a child it spawns sees no token.
5085+
- **stubbed:** no live run against a real Hub or a real Electron window; the "Install Lucid Meeting Hub" row carries no install URL because there is no published one to point at; talk-time bar and `/ext/analytics` are not consumed; mark-done is one-way (the open list is all the panel can see, so un-doing stays a Hub action).
5086+
- **next:** Nick's three calls from LUCIDMeetingHub#7 - left-rail placement/priority, pairing UX now vs Plan 8 device certs, and whether mark-done belongs in v1 at all; then pair against a live Hub, walk dormant -> unpaired -> locked -> unlocked in the running app, and rebuild `renderer/app.bundle.js` before believing anything on screen (ADR-0260).
50775087

50785088
## P-LEGIBLE.1: legible to Defender and Agent 365 without a content path (ADR-0384, issue #302)
50795089
- **shipped:** investigated Defender local-agent discovery and runtime protection against Microsoft's primary docs: discovery is a Microsoft-maintained supported list (no public registration), Entra Agent ID does not apply to local agents, runtime protection rides vendor hook interfaces, and network inspection cannot see pinned/HTTP/3/loopback traffic. New pure `desktop/local_agent_manifest.ts` + a boot write in `dev.ts` (only when Electron launched it; `main.ts` passes `LUCID_HOST_EXE`): each launch writes `<userData>/local-agent-manifest.json` (standard build `%APPDATA%\lucidagentide-desktop`, Creator `LucidCreator`), an advisory LUCID-defined file whose field names borrow Defender's profile vocabulary (vendor, version, relatedProcess, processes, autoApprove as a string, mcpServers, localMcps) plus our posture (loopback + ADR-0024 token, no agent-native hook, network inspection not effective); Microsoft publishes no vendor-writable manifest format, so writing it enrolls nothing. Metadata only by construction: MCP entries keep name, type, URL origin or command basename, never headers, args, env, path or query. Review fixes: the runbook and Intune/live-response commands now name the real userData folder (`lucidagentide-desktop`, not `LucidAgentIDE`); new `desktop/build/installer.nsh` (`build.nsis.include`) deletes only the manifest on a real NSIS uninstall so it cannot report a removed install, and the documented detection check requires the installed `LucidAgentIDE.exe` from the uninstall entry. Admin runbook `docs/DEFENDER-AGENT365-COEXISTENCE.md` with the CUI posture ("network inspection only") and collection/hunting recipes. `make demo-P-LEGIBLE.1`.

0 commit comments

Comments
 (0)