Repository navigation
docs(m17): 17.2a Linux VM spike — measure linux-x86_64, label the boundary - #84
Conversation
…undary Split 17.2 into 17.2a (Linux VM, done) and 17.2b (real hardware, blocked on procurement). None of the hardware D-M17-3 names exists, and closing 17.2 with a VM run would have redefined "done" by omission in the one milestone whose stated binding constraint is verification. Measured on WSL2 Ubuntu 24.04.4, tier VM-verified -- NOT hardware-verified: - capture end-to-end on POSIX (3 raw / 16 events / 1 machine in an isolated archive) - 17.1's --home confinement: 0 declared paths outside the given home, 0 naming /mnt/c - the read-side check 16.0's write-only checks missed: 1120 /proc/<pid>/fd samples - tsc -b --force 9.47 s; vitest 1045 | 680 (= 1034 + 17.1's net +11 tests, and 1045 + 680 = the real repo's 1725 | 0) - quickstart breaks in the SAME two places as Windows with byte-identical errors, so INC-2026-02/03 are onboarding defects, not Windows quirks - SIGINT graceful drain 503 ms, clean exit Six incidents, INC-2026-08..13 (0 high, 5 med, 1 low): - 09: build-sea.mjs exits 0 printing PASS while emitting a working 118 MB ELF named collector-x86_64-pc-windows-msvc.exe that Tauri's externalBin cannot find - 08: build-sea.mjs --check boots serve with no --home, creating the operator's real ~/.420ai -- a double-writer hazard, and the one isolation constraint this run breached. Repaired in place under D-16.0-2 and recorded, not hidden - 10: npm run setup aims DATABASE_URL_TEST at whatever cluster is already on the host, so a second checkout's vitest run would TRUNCATE the real test DB (demonstrated, not fired) - 11: INC-2026-05's recorded root cause contradicts file-watcher.ts:136, unchanged since 2026-06-13 -- a static session file IS captured, and 16.3 was designed against the wrong cause - 12: a connector source path that does not exist leaves no trace anywhere - 13: WSL interop fabricated a plausible wrong value in the transcript 17.3 re-scored 65% -> 80%: the whole SEA recipe runs unmodified on Linux, so the risk was the mechanism and the mechanism works. 17.4/17.5/17.6 deliberately NOT re-scored, with the reason stated per row -- systemd is off in default WSL2, webkit2gtk-4.1 is missing, and installers/Gatekeeper need hardware. Reading this run as "Linux OK" would be the emulated != hardware-verified failure the slice was shaped to prevent. Introduces the verification-tier vocabulary (hardware-verified / VM-verified / CI-verified / not measured) as a required field on findings, generalizing D-M17-4. No product code changed (D-M17-5). Docs-only diff. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ource edit Adds the code review and execution report, and fixes all four review findings (maintainer chose "fix everything" at the triage gate). Report accuracy, all in the 17.2a clean-room report: - "1120 fd samples" conflated sampling ROUNDS with path OBSERVATIONS. The loop is 160 rounds at 4 Hz over 40 s and each round records every open path, so 1120 is 7 observations x 160 rounds -- the original phrasing overstated temporal coverage 7x, in the one number the read-side evidence rests on. Restated at all four sites (report x2, milestone plan, SUMMARY). The correction makes the existing "a read-and-close inside one 250 ms window can be missed" caveat MORE load-bearing, which is the honest direction. - Scoped the headline row to "Real WINDOWS credentials / archive modified". It was accurate but unscoped, and the run did create a real WSL-side collector home (INC-2026-08, stated two rows down) -- the only row that could be quoted misleadingly. - New section 4b derives the capture counts from their inputs (fixture, 4 seeded files, seed/append order, file_cursors=4) and DECLARES raw=3 unexamined rather than leaving a reader to infer a capture gap. It was the one result in the report presented without its derivation. The one source edit, disclosed because it changes what the diff is: - apps/collector/src/connectors/cursor-store.ts jsdoc addressed its measurement debt to "17.2" -- a slice number THIS BRANCH split into 17.2a (done) and 17.2b (blocked). A reader would see "17.2", find 17.2a marked done, and conclude the debt is discharged: the exact misreading the report warns about, at the one place the decision gets made. Now says 17.2b, and separates CONFINEMENT (verified by 17.2a on linux-x86_64) from CONVENTION (still unmeasured on darwin and linux). Comment-only. No executable product code changed and no defect this slice MEASURED was fixed -- INC-2026-08..12 all stand (D-M17-5). Verified the edit cannot move the §10.4 capture-surface fingerprint: captureSurfaceFingerprint hashes watchGlobs, requiredPermissions, poll.sources and push.origins -- runtime values, never source text -- so no install flips to needs-approval (the D-17.1-3 hazard, checked). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… last fix Re-review pass 2 caught a defect in the previous commit's own fix, which is what the code-review -> fix -> code-review loop is for. Section 4b said the seeded session file "was later appended", AFTER the three extra files were written. The real order is the reverse: append at step 3, extra files at step 4. The order is load-bearing -- step 2 -> 3 is the INC-2026-05 evidence (a static file captured with no append), so getting it backwards undermines the finding it exists to support. Rewritten as a stepwise table carrying the archive counts observed at each point. That also exposes something the prose had hidden: step 4 added 3 files and 6 lines, moved events by +6, and moved raw not at all. Now declared unexamined and flagged for 17.2b rather than left invisible. Pass 3 is clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Documents and records M17 slice 17.2a (Linux VM / WSL2) as a measurement-only spike, explicitly labeling what was and wasn’t measurable on that substrate, and splitting the original 17.2 hardware spike into 17.2a (VM-verified) and 17.2b (hardware-verified, blocked on procurement) across the repo’s planning + incident tracking artifacts.
Changes:
- Update milestone tracking (
SUMMARY.md) to reflect the 17.2a/17.2b split, mark 17.2a done, and clarify re-scoring boundaries. - Add the 17.2a artifacts (plan + clean-room report + execution report + code review) and extend the incident log with INC-2026-08…13 including a required “verification tier” label.
- Adjust
cursor-store.tsjsdoc to reference 17.2b for hardware verification, and to explicitly separate “confinement verified” vs “convention unmeasured”.
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| SUMMARY.md | Marks 17.2a as complete, introduces 17.2b as procurement-blocked, and records what was (and wasn’t) re-scored by the VM spike. |
| apps/collector/src/connectors/cursor-store.ts | Updates jsdoc to point the “hardware-verified” debt at 17.2b and clarify what 17.2a did/didn’t verify. |
| .agents/research/incidents.md | Adds verification-tier policy + known-defect rule guidance and logs new incidents INC-2026-08…13. |
| .agents/research/cleanroom-linux-x86_64-2026-08-09.md | New clean-room measurement report for linux-x86_64 on WSL2 (VM-verified), including explicit “not measured” boundary. |
| .agents/plans/m17-slice2a-linux-vm-spike.md | New plan defining scope, methodology, and acceptance criteria for the 17.2a spike. |
| .agents/plans/m17-cross-platform-collectors.md | Updates the milestone plan to include the 17.2a/17.2b split, adds measured results for 17.2a, and re-scores 17.3. |
| .agents/execution-reports/m17-slice2a-linux-vm-spike.md | New execution report capturing what ran, what diverged, and validation outcomes for the 17.2a spike. |
| .agents/code-reviews/m17-slice2a-linux-vm-spike.md | New code-review artifact documenting review checks and resolved findings for the spike’s artifacts. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| @@ -0,0 +1,630 @@ | |||
| # Feature: M17 slice 17.2a — Linux VM spike (the measurable half of 17.2) | |||
|
|
|||
| The following plan should be complete, but its important that you validate documentation and codebase | |||
…ot wrong Seven specialist agents reviewed PR #84. The valuable finding is that a slice whose whole thesis is "an unverified claim is a defect" shipped several. All fixed. Wrong claims that would have misdirected later slices: - INC-2026-08's root cause cited cli.ts:631 resolveHome, but `serve` is its own entrypoint and never consults it: cli.ts dispatches no `serve` command, serve.ts:572 only tests argv.includes("serve") and parses no flags, and the home is pinned in THREE places (serve.ts:107 QUEUE_PATH import-time constant; runEngine called with neither home nor queuePath so capture-engine.ts:363-364 falls back; loadCredentials -> CREDENTIALS_PATH). So the prescribed "one --home flag on two call sites" fix would be silently ignored and the prescribed regression test would FAIL with the fix applied. Rewritten: serve defaulting to homedir() is correct for its real job (it is the desktop/service sidecar, serve.ts:291) -- the defect is that a BUILD SCRIPT invokes it, so the fix belongs in the caller (env override), and the test must assert argv/env or use a hermetic HOME override. - The 17.3 re-score was over-generalized. 17.3 is MULTI-target; the run measured ELF only. build-sea.mjs has no --macho-segment-name NODE_SEA and no codesign remove/re-sign (grep -c 'macho-segment-name|codesign|darwin' -> 0), both of which Node's SEA recipe requires on macOS, and ad-hoc signing intersects D-M17-2. Re-scored to "80% for the LINUX triples; darwin unchanged at 65%, tier not measured". - scripts/sidecar-stub.mjs was never run (a protocol step) and never recorded as skipped -- and it already exports sidecarFileName(hostTriple, platform) with a `rustc -vV` derivation, which is the strongest available support for the 17.3 re-score. Recorded, and cited in INC-2026-09. - "1045 + 680 = 1725" was a tautology sold as corroboration: vitest prints the total in the same line, so passed+skipped==total always. Replaced with the check that can fail -- all 63 skipIf sites gate on DATABASE_URL_TEST, no test is platform-gated, and the 61 skipped files are exactly the *.int.test.ts set. - Three regression-test lines could not fail: INC-2026-08's "real home untouched" (this report's own baseline section explains why -- and its negative control was DESTRUCTIVE), INC-2026-09's self-referential basename oracle with an inert control, and INC-2026-10's derivation test whose oracle has no referent. All rewritten with controls that discriminate; both build-sea tests also record an unacknowledged precondition (the script has zero exports). - Two wrong incident IDs in the report (INC-2026-10 where INC-2026-11 was meant), the pre-sign-off tally ("five routed to 17.3/17.4" -> three), 17.4's "measured nothing here" (INC-2026-12 was measured by 17.2a and routed to 17.4), and incidents.md's preamble ("the five below" -> six; INC-2026-01 IS fixed, in 17.1). - D-M17-3 still read "all hardware-verified ... a Linux VM", asserting exactly what this slice exists to deny. Amended with its true current state. - npm ci was never run (the documented path says npm install); the lockfile-integrity half of the named residual risk is not measured, not resolved. - Row 10 (case-sensitive dedup) was routed to 17.2b, which is procurement-blocked -- but the test needs no hardware, so parking it there means it never happens. Re-routed to 17.3/17.7 with the other substrate-independent leftovers. - Level 4 was claimed "all 7 boxes, checked mechanically"; box 3 (every results row carries a tier) failed on four tables. Fixed and disclosed. - 17.7 was silently unaddressed though its premise is now known to be procurement-blocked; it also now owns enforcing the tier field. Source (comment-only, maintainer's call at the triage gate): - cursor-store.ts:35-36 documented the poll loop as reporting unavailable connector health -- which THIS SLICE measured as false (INC-2026-12), while :105-106 of the same file already said the opposite. Corrected, naming the mechanism. - The new jsdoc claim carried NO verification tier, violating the acceptance criterion this branch itself added. Now tiered, and it separates purity from confinement (a constant is pure and confines nothing) and names the guarding tests. - connector.ts documented PollOutcome.unavailable as surfaced/degraded; it is written but never read. D-M17-5 gains a written second exception (comment-only correction of a reference THIS SLICE made stale), bounded and recorded where the decision lives rather than only in an execution report. npm run repo-health -- --require-db: PASS, 680 integration tests ran, 0 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
prp-review (
|
| # | Severity | Finding | File | Disposition | What was done | Status |
|---|---|---|---|---|---|---|
| 1 | Critical | INC-2026-08's root cause cites cli.ts:631 resolveHome, but serve never consults it — its prescribed fix would be silently ignored and its regression test would fail with the fix applied |
.agents/research/incidents.md |
Fix | Rewrote root cause (3 home pins: serve.ts:107, capture-engine.ts:363-364, identity.ts:20), disposition (env override in the caller; serve's own default is correct for its job) and test |
Fixed |
| 2 | Important | 17.3 re-score over-generalized — 17.3 is multi-target; only ELF was measured. No --macho-segment-name, no codesign |
m17-cross-platform-collectors.md:181 |
Fix | Re-scored to 80% Linux triples / 65% darwin, tier not measured; Mach-O SEA carried to 17.2b |
Fixed |
| 3 | Important | 1045 + 680 = 1725 is a tautology — vitest prints the total; passed + skipped == total always |
cleanroom-…-08-09.md:448 |
Fix | Replaced with the check that can fail: all 63 skipIf gate on DATABASE_URL_TEST, none platform-gated, 61 skipped files == the *.int.test.ts set |
Fixed |
| 4 | Important | 3 of 6 regression-test lines cannot fail — and INC-2026-08's negative control was destructive (would boot a second collector against the live queue) | incidents.md (08, 09, 10) |
Fix | All three rewritten with discriminating controls; both build-sea tests now record the unacknowledged precondition (the script has zero exports) |
Fixed |
| 5 | Important | scripts/sidecar-stub.mjs — a prescribed protocol step — was never run and never recorded as skipped; it already ships sidecarFileName(hostTriple, platform) |
report · incidents.md |
Fix | Recorded as Methodology correction #5; cited in INC-2026-09 as the reuse target for 17.3 | Fixed |
| 6 | Important | Two wrong incident IDs (INC-2026-10 where 11 was meant) |
cleanroom-…-08-09.md:171,361 |
Fix | Corrected; found by 4 of 7 agents | Fixed |
| 7 | Important | Pre-sign-off tally "five routed to 17.3/17.4" → actually three | m17-cross-platform-collectors.md:327 |
Fix | Corrected with per-ID routing | Fixed |
| 8 | Important | incidents.md preamble: "the five below" → six; "none is fixed" → INC-2026-01 shipped fixed in 17.1 |
incidents.md:70-74 |
Fix | Corrected, with a note on how long it stood | Fixed |
| 9 | Important | 17.4's "17.2a measured nothing here" contradicts INC-2026-12 — measured by 17.2a, routed to 17.4 | m17-cross-platform-collectors.md:180 |
Fix | Split into service-manager half (nothing measured) vs INC-2026-12 (measured, insufficient to move the score) | Fixed |
| 10 | Important | D-M17-3 still read "all hardware-verified … a Linux VM" — asserting exactly what this slice exists to deny | m17-cross-platform-collectors.md:230 |
Fix | Amended with true current state: none of the three exists | Fixed |
| 11 | Important | npm ci never ran (documented path says npm install); the lockfile half of the named risk was declared resolved |
cleanroom-…-08-09.md:233 |
Fix | §2 retitled; lockfile-integrity half marked not measured |
Fixed |
| 12 | Important | Row 10 (case-sensitive dedup) parked in procurement-blocked 17.2b though it needs no hardware | m17-cross-platform-collectors.md:107 |
Fix | Re-routed to 17.3/17.7 with the other substrate-independent leftovers | Fixed |
| 13 | Important | "Level 4 ✓ all 7 boxes, checked mechanically" — box 3 fails on 4 tables | exec report :47 |
Fix | Per-table tier statements added; the overclaim disclosed | Fixed |
| 14 | Important | 17.7 silently unaddressed though its premise is now procurement-blocked | m17-cross-platform-collectors.md:183 |
Fix | Marked not-re-scored with reason; now owns tier-field enforcement | Fixed |
| 15 | Important | cursor-store.ts:35-36 documents the poll loop as reporting unavailable health — this slice measured that false; :105-106 says the opposite |
apps/collector/src/connectors/cursor-store.ts |
Fix | Corrected, naming the mechanism (INC-2026-12) | Fixed |
| 16 | Important | The new jsdoc claim carried no verification tier — violating the criterion this branch added | cursor-store.ts:12 |
Fix | Tiered; also separates purity from confinement and names the guarding tests | Fixed |
| 17 | Suggestion | PollOutcome.unavailable documented as "surfaced"/"degraded"; written but never read |
connector.ts:101,112 |
Fix | Corrected | Fixed |
| 18 | Suggestion | Stale refs: 9 bare 17.2, incidents.md:150 (off by 251 lines), cursor-store.ts:11-13 → :11-14 |
plan · SUMMARY · report | Fix | Swept; genuinely historical references left intact | Fixed |
| 19 | Suggestion | SUMMARY §0 entry 2× its neighbours, near-duplicating §6; re-score prose in 3 places; 17.2b list 5× | SUMMARY.md |
Fix | §0 collapsed to the 17.0/17.1 register; plan's slice table made the single confidence ledger | Fixed |
| 20 | Suggestion | Verification tier declared required but absent from the entry-format template |
incidents.md:47 |
Fix | Added to the template — the mechanism that makes it survive the next author | Fixed |
| 21 | Suggestion | INC-2026-11 cited one line for behaviour holding across all three capture modes | incidents.md |
Fix | Snapshot + poll evidence added ("first sight = full ingest in every mode") | Fixed |
| 22 | Suggestion | INC-2026-08 tiered VM-verified, but the defect is platform-independent — only detectability was Linux-specific |
incidents.md:288 |
Fix | Tier line qualified with the sharper point | Fixed |
| 23 | Suggestion | §4b's event fan-out undeclared (2 lines → 8 events vs 6 lines → 6) | cleanroom-…-08-09.md |
Fix | Both anomalies now explicitly declared unexamined | Fixed |
| 24 | Suggestion | "the CONVENTION below" — the table is above | cursor-store.ts:14 |
Fix | Corrected | Fixed |
| 25 | Suggestion | Code review's "docs-only" claim not marked superseded; its consistency check never covered incident cross-refs | .agents/code-reviews/… |
Fix | Both recorded, including the gap that let finding 6 through | Fixed |
Verified, not accepted on trust. Every Critical/Important was re-derived against source before acting: cli.ts dispatches no serve command; serve.ts:572 parses no flags; grep -c for Mach-O/codesign handling in build-sea.mjs returns 0; sidecar-stub.mjs:50 exports the helper; and captureSurfaceFingerprint hashes runtime values only, so no install flips to needs-approval.
Not fixed, out of scope, flagged deliberately: cleanroom-2026-08-02.md:147 contains a full ingest token in cleartext — a real §3 privacy violation predating this PR (16.0, throwaway DB since dropped). Belongs in its own commit.
Also noted: /lril:prp-review's --agents mode references a workflows/agents.md that is not installed — only the flat command file exists. The seven aspects were derived from the command file's own mode-select table.
Summary: 25 findings — 25 fixed, 0 declined, 0 deferred. D-M17-5 gains a written second exception (comment-only correction of a reference this slice made stale), bounded and recorded where the decision lives rather than only in an execution report. npm run repo-health -- --require-db — PASS, 680 integration tests ran, 0 skipped. Dispositions were chosen by the maintainer at the triage gate.
What
Slice 17.2a executes the subset of the 17.2 spike protocol that an available substrate can measure truthfully, and is explicit about the far larger subset it cannot.
17.2 as written is a clean-room spike on three hardware targets (D-M17-3). None of that hardware exists. Rather than quietly redefine "17.2 is done", this ships as 17.2a (Linux VM, tier
VM-verified) with a new 17.2b row carrying everything the boundary excludes, visibly blocked on procurement. M17's own framing — "verification, not implementation, is the binding constraint" — makes that the one substitution this milestone must never make.The headline names what was NOT measured before what was. Slices 17.4/17.5/17.6 — the three lowest-confidence slices, i.e. precisely the ones 17.2 exists to re-score — are not re-scored, because a VM cannot see systemd,
webkit2gtk, a desktop session, an installer or Gatekeeper.Measured (WSL2 Ubuntu 24.04.4, real kernel + ext4 + POSIX
homedir())--homeconfinement holds — 0 declared paths outside the given home, 0 naming/mnt/c/proc/<pid>/fdsampling rounds; only clean-room files ever opentsc -b --force9.47 s;vitest1045 | 680 (= 1034 + 17.1's net +11 tests; 1045 + 680 = the real repo's 1725 | 0)quickstart.mdbreaks in the SAME two places as Windows, byte-identical errors — so INC-2026-02/03 are onboarding defects, not Windows quirkssystem_identifier, not by assertion (INC-2026-06's prescribed containment, verified working)Incidents raised — INC-2026-08…13 (0 high, 5 med, 1 low)
build-sea.mjsexits 0 printingPASSwhile emitting a working 118 MB ELF namedcollector-x86_64-pc-windows-msvc.exethat Tauri'sexternalBincannot findbuild-sea.mjs --checkbootsservewith no--home, creating the operator's real~/.420ai: a double-writer hazard, and the one isolation constraint this run breached. Repaired in place under D-16.0-2 and recorded, not hiddennpm run setupaimsDATABASE_URL_TESTat whatever cluster is already on the host, so a second checkout'snpx vitest runwould TRUNCATE the real test DB (demonstrated, not fired)file-watcher.ts:136, unchanged since 2026-06-13. A static session file is captured — and 16.3's scorecard was designed against the wrong causepoll_state=0,connector_errors=0, no log line)Decisions
hardware-verified/VM-verified/CI-verified/not measured), generalizing D-M17-4 —emulated ≠ hardware-verifiedis the next member of theskipped ≠ passedfamily.cursor-store.tsjsdoc addressed its debt to "17.2", a number this branch split) was made at the maintainer's triage decision and is disclosed in the execution report. Verified it cannot move the §10.4 fingerprint.Review
.agents/code-reviews/m17-slice2a-linux-vm-spike.md— 4 findings, all resolved; pass 2 caught a defect in its own fix, pass 3 clean.agents/execution-reports/m17-slice2a-linux-vm-spike.md.agents/research/cleanroom-linux-x86_64-2026-08-09.mdValidation
npm run repo-health -- --require-db— PASS, 680 integration tests ran, 0 skippednpm run lint/npm run format:check— PASSbuild:dashboard— not run;apps/dashboarduntouched by this slicerepo-healthwas run three times during the slice and the middle run failed (ECONNREFUSED, Docker Desktop exited and took the archive down mid-crash-recovery — not a leak; INC-2026-06's signature is28P01). Disclosed in the report.🤖 Generated with Claude Code