Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 30 additions & 2 deletions devflow-plugin/commands/verify-first.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
description: [0.27.3] Ground-truth every technical claim — and every requirement-coverage claim — in a plan or design against the source of truth (run the code, hit the API with curl, probe the live browser DOM/store, build a proof-of-concept; for requirements, diff the plan section-by-section against the original design doc + ticket) before it can be relied on. Reading docs, blogs, training knowledge, or trusting that the plan captured the requirements is NOT verification. Use when stress-testing assumptions, before locking an architectural decision, or when the user says "verify first", "verify with practical tests", "prove it", or "test, don't assume".
description: [0.27.3] Ground-truth every technical claim — and every requirement-coverage claim — in a plan or design against the source of truth (run the code, hit the API with curl, probe the live browser DOM/store, build a proof-of-concept; for requirements, diff the plan section-by-section against the original design doc + ticket) before it can be relied on. Reading docs, blogs, training knowledge, or trusting that the plan captured the requirements is NOT verification, and neither is a probe that could not have failed (an HTTP 200, a bound port, a detector that reported nothing because it never ran). Use when stress-testing assumptions, before locking an architectural decision, or when the user says "verify first", "verify with practical tests", "prove it", or "test, don't assume".
---

# Verify First
Expand Down Expand Up @@ -33,6 +33,27 @@ require('CometRelayEnvironment').getStore().getSource().getRecordIDs()
1. A negative result on ONE probe is **not** proof of absence. `window.Relay` being undefined said nothing about `CometRelayEnvironment`.
2. **Never refute a claim you have not actually tested.** "It doesn't do X" is itself a claim that needs its own positive test.

## The inverse failure: a probe that passes for the wrong reason

The Relay case is a false NEGATIVE — the probe said "absent" while the thing was present. The mirror image costs just as much and is far easier to miss, because it feels like success: the probe passes, you write `[V]`, and the thing you actually cared about was broken the whole time.

A skill whose entire job was to hand a reviewer a running app was recorded as "fully verified end to end" on the strength of `HTTP 200` plus a `<title>` tag. The app was broken in the browser — `does not provide an export named 'AIAgent'` — because a stale generated module never resolved. A vite dev server returns 200 for the HTML shell even when the ES modules it references fail, so that probe **could not have failed** for the reason that mattered. A human opening the page found it in seconds.

**Liveness is not correctness.** Each of these proves something started, never that it works:

| Signal | What it proves | What it does NOT prove |
|---|---|---|
| HTTP 2xx | something answered the socket | the response is right, or the app renders |
| A page title / HTML shell | the server emitted markup | the JS resolved, hydrated, or ran |
| A live PID / a bound port | a process started | it is serving what you just built |
| Exit code 0 | the command ran | it did the work — many tools exit 0 on a no-op |
| A startup log banner | execution reached that line | startup finished, or survived |
| "0 problems found" | the reporter printed | the checker ever ran |

**Detectors fail silent-clean.** For anything whose output is a count of problems — a linter, an a11y scan, a console-error capture, a filtered test run — "found nothing" and "never ran" are byte-identical. A green detector is evidence only after you have watched it go RED on a known-bad input. Keep a deliberately broken fixture and prove each check FIRES before you trust its silence.

**Lesson 3, alongside the two above: match the precision of the claim to the precision of the probe.** If all you polled was the port, then the verified claim is "the server is up" — write exactly that, not "the app works". Quietly widening what a probe established is how a `[B]` gets laundered into `[V]` without anyone noticing.

## Workflow

Turn this checklist into TodoWrite items at the start (one todo per step, plus one todo per `[B]` claim once enumerated). Do not skip a step because the claim "looks obvious".
Expand Down Expand Up @@ -87,7 +108,9 @@ For each `[B]` claim, prefer the most direct real-system probe available:
- A REPL or a throwaway unit spike over reasoning about the type signatures.
- A shell one-liner against the real binary / file over assuming flag behaviour.

Capture the ACTUAL output as evidence (the record count, the response body, the exit code, the thrown error). Paste it into the ledger. If a probe comes back negative, do not conclude absence — run a second, differently-shaped probe before refuting the claim (lesson 1 above).
**Choose a probe that can actually fail for the reason you care about.** Before running it, ask: *if the thing I am checking were broken, would this probe look any different?* If the answer is no, it is the wrong probe and passing it proves nothing (lesson 3 above). A readiness poll cannot see a broken import. A type-check cannot see wrong behaviour. A green build cannot see a runtime crash. Move the probe as close as you can to the failure you are actually worried about: render the page and read the console rather than reading the status line.

Capture the ACTUAL output as evidence (the record count, the response body, the exit code, the thrown error). Paste it into the ledger. If a probe comes back negative, do not conclude absence — run a second, differently-shaped probe before refuting the claim (lesson 1 above). If it comes back positive, check that it was *capable* of coming back negative.

## Step 4 — Record the verification ledger

Expand Down Expand Up @@ -138,6 +161,11 @@ A decision may NOT be locked, written into the plan as settled, or built upon wh
| "I already know this lib does X" | Knowledge rots. Run the probe. The user's override beats your instinct. |
| "The docs say it returns Y, that's enough" | A doc/blog citation is a `[B]` claim, not `[V]`. A link is not evidence — curl it / stream it. |
| "`window.X` is undefined, so the feature is gone" | One negative probe ≠ absence. Try a structurally different probe before refuting. |
| "It returned 200, so it works" | 200 is liveness, not correctness — a dev server serves the shell even when every module fails. Probe the thing you actually care about. |
| "The scan reported 0 violations" | Identical output to a scan that never ran. Watch it go red on a known-bad fixture before trusting its silence. |
| "The process is up / the port is bound" | That is startup, not function. Ask it to do the work and check the answer. |
| "It exited 0, so it did the thing" | Plenty of tools exit 0 on a silent no-op or a skipped run. Assert the effect, never the exit code alone. |
| "My probe passed, so the claim is `[V]`" | Only if it could have failed. Match the claim to what the probe actually established — "the server is up" is not "the app works". |
| "I'll just write my own debounce/fuzzy-match, it's tiny" | Solved problem = untested edge-case claim. Pick a battle-tested lib, log it in the inventory (Step 5). |
| "I'll match on the 'Sold'/'Sponsored' label text" | Locale-fragile — breaks on translation. Anchor on DOM structure / class / id / role instead. |
| "I can lock this decision, the test is trivial" | Trivial means run it now. `[B]` is a blocker — never ship it unrun. |
Expand Down
32 changes: 30 additions & 2 deletions devflow-plugin/skills/verify-first/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: verify-first
description: Ground-truth every technical claim — and every requirement-coverage claim — in a plan or design against the source of truth (run the code, hit the API with curl, probe the live browser DOM/store, build a proof-of-concept; for requirements, diff the plan section-by-section against the original design doc + ticket) before it can be relied on. Reading docs, blogs, training knowledge, or trusting that the plan captured the requirements is NOT verification. Use when stress-testing assumptions, before locking an architectural decision, or when the user says "verify first", "verify with practical tests", "prove it", or "test, don't assume".
description: Ground-truth every technical claim — and every requirement-coverage claim — in a plan or design against the source of truth (run the code, hit the API with curl, probe the live browser DOM/store, build a proof-of-concept; for requirements, diff the plan section-by-section against the original design doc + ticket) before it can be relied on. Reading docs, blogs, training knowledge, or trusting that the plan captured the requirements is NOT verification, and neither is a probe that could not have failed (an HTTP 200, a bound port, a detector that reported nothing because it never ran). Use when stress-testing assumptions, before locking an architectural decision, or when the user says "verify first", "verify with practical tests", "prove it", or "test, don't assume".
---

# Verify First
Expand Down Expand Up @@ -34,6 +34,27 @@ require('CometRelayEnvironment').getStore().getSource().getRecordIDs()
1. A negative result on ONE probe is **not** proof of absence. `window.Relay` being undefined said nothing about `CometRelayEnvironment`.
2. **Never refute a claim you have not actually tested.** "It doesn't do X" is itself a claim that needs its own positive test.

## The inverse failure: a probe that passes for the wrong reason

The Relay case is a false NEGATIVE — the probe said "absent" while the thing was present. The mirror image costs just as much and is far easier to miss, because it feels like success: the probe passes, you write `[V]`, and the thing you actually cared about was broken the whole time.

A skill whose entire job was to hand a reviewer a running app was recorded as "fully verified end to end" on the strength of `HTTP 200` plus a `<title>` tag. The app was broken in the browser — `does not provide an export named 'AIAgent'` — because a stale generated module never resolved. A vite dev server returns 200 for the HTML shell even when the ES modules it references fail, so that probe **could not have failed** for the reason that mattered. A human opening the page found it in seconds.

**Liveness is not correctness.** Each of these proves something started, never that it works:

| Signal | What it proves | What it does NOT prove |
|---|---|---|
| HTTP 2xx | something answered the socket | the response is right, or the app renders |
| A page title / HTML shell | the server emitted markup | the JS resolved, hydrated, or ran |
| A live PID / a bound port | a process started | it is serving what you just built |
| Exit code 0 | the command ran | it did the work — many tools exit 0 on a no-op |
| A startup log banner | execution reached that line | startup finished, or survived |
| "0 problems found" | the reporter printed | the checker ever ran |

**Detectors fail silent-clean.** For anything whose output is a count of problems — a linter, an a11y scan, a console-error capture, a filtered test run — "found nothing" and "never ran" are byte-identical. A green detector is evidence only after you have watched it go RED on a known-bad input. Keep a deliberately broken fixture and prove each check FIRES before you trust its silence.

**Lesson 3, alongside the two above: match the precision of the claim to the precision of the probe.** If all you polled was the port, then the verified claim is "the server is up" — write exactly that, not "the app works". Quietly widening what a probe established is how a `[B]` gets laundered into `[V]` without anyone noticing.

## Workflow

Turn this checklist into TodoWrite items at the start (one todo per step, plus one todo per `[B]` claim once enumerated). Do not skip a step because the claim "looks obvious".
Expand Down Expand Up @@ -88,7 +109,9 @@ For each `[B]` claim, prefer the most direct real-system probe available:
- A REPL or a throwaway unit spike over reasoning about the type signatures.
- A shell one-liner against the real binary / file over assuming flag behaviour.

Capture the ACTUAL output as evidence (the record count, the response body, the exit code, the thrown error). Paste it into the ledger. If a probe comes back negative, do not conclude absence — run a second, differently-shaped probe before refuting the claim (lesson 1 above).
**Choose a probe that can actually fail for the reason you care about.** Before running it, ask: *if the thing I am checking were broken, would this probe look any different?* If the answer is no, it is the wrong probe and passing it proves nothing (lesson 3 above). A readiness poll cannot see a broken import. A type-check cannot see wrong behaviour. A green build cannot see a runtime crash. Move the probe as close as you can to the failure you are actually worried about: render the page and read the console rather than reading the status line.

Capture the ACTUAL output as evidence (the record count, the response body, the exit code, the thrown error). Paste it into the ledger. If a probe comes back negative, do not conclude absence — run a second, differently-shaped probe before refuting the claim (lesson 1 above). If it comes back positive, check that it was *capable* of coming back negative.

## Step 4 — Record the verification ledger

Expand Down Expand Up @@ -139,6 +162,11 @@ A decision may NOT be locked, written into the plan as settled, or built upon wh
| "I already know this lib does X" | Knowledge rots. Run the probe. The user's override beats your instinct. |
| "The docs say it returns Y, that's enough" | A doc/blog citation is a `[B]` claim, not `[V]`. A link is not evidence — curl it / stream it. |
| "`window.X` is undefined, so the feature is gone" | One negative probe ≠ absence. Try a structurally different probe before refuting. |
| "It returned 200, so it works" | 200 is liveness, not correctness — a dev server serves the shell even when every module fails. Probe the thing you actually care about. |
| "The scan reported 0 violations" | Identical output to a scan that never ran. Watch it go red on a known-bad fixture before trusting its silence. |
| "The process is up / the port is bound" | That is startup, not function. Ask it to do the work and check the answer. |
| "It exited 0, so it did the thing" | Plenty of tools exit 0 on a silent no-op or a skipped run. Assert the effect, never the exit code alone. |
| "My probe passed, so the claim is `[V]`" | Only if it could have failed. Match the claim to what the probe actually established — "the server is up" is not "the app works". |
| "I'll just write my own debounce/fuzzy-match, it's tiny" | Solved problem = untested edge-case claim. Pick a battle-tested lib, log it in the inventory (Step 5). |
| "I'll match on the 'Sold'/'Sponsored' label text" | Locale-fragile — breaks on translation. Anchor on DOM structure / class / id / role instead. |
| "I can lock this decision, the test is trivial" | Trivial means run it now. `[B]` is a blocker — never ship it unrun. |
Expand Down
Loading
Loading