This document describes how Spectre's pieces fit together. The user-facing surface (/vision, /implement, the hooks) is documented in README.md; this is the internals view.
┌────────────────────────────────────────────────────────────┐
│ CONTEXT PLANE (hooks — own what Claude Code "sees") │
│ SessionStart → bin/hydrate.py │
│ PostToolUse(Bash) → bin/compact.py │
└────────────────────────────────────────────────────────────┘
┌────────────────────────────────────────────────────────────┐
│ AGENT PLANE (skills — own the multi-turn protocol) │
│ /vision → skills/vision/SKILL.md │
│ /implement → skills/implement/SKILL.md │
└────────────────────────────────────────────────────────────┘
┌────────────────────────────────────────────────────────────┐
│ STATE PLANE (bin/ — own all on-disk truth) │
│ Specs: specs/.active, specs/<slug>.spec.md │
│ Graph: specs/.graph.md │
│ Scratchpad: state/scratchpad.json (v2, per-track) │
│ Bundle: state/.eval-bundle.json (transient) │
│ Sidecar: <spec>.eval.json (post-lock) │
│ ADRs: decisions/<NNNN>-<slug>.md │
│ Locks: runtime/supervisor.{sock,pid} │
└────────────────────────────────────────────────────────────┘
The strict separation matters. Hooks must finish in <10s and emit only additionalContext; they never block, never prompt, never write specs. Skills are multi-turn, prompt for confirmation, and are the only writers of .active. Modules under bin/ are pure functions over disk state — they have no opinion about what Claude Code shows the user.
- Detect v1 scratchpad and migrate to v2 in place (idempotent —
migrate_scratchpad_v1_to_v2.py). - Read
specs/.active. If absent: emitSIGNAL: No active spec. Run /vision to begin.plus the list of available specs. - If present: read the spec body and emit it wrapped in
--- ACTIVE SPEC ---markers, followed by a per-trackSTATE: step=N exit_code=X last_command=…line. - Stale pointer (file referenced by
.activeno longer exists): emitERROR: stale .active pointer ….
- Parse the Bash result into a Delta string via
parse_delta_with_paths(cmd)— regex-driven, captures both the human-readable delta (pip install,mkdir foo) and the affected paths. - Update
state/scratchpad.json:last_command,exit_code,delta,timestampalways overwritten.- On success (
exit_code == 0): dedupe-append the captured paths topaths_touched(capped at 200 FIFO). - On failure: append a
failed_hypotheses[]entry with the first matching error line. Never softened.
- Emit
additionalContextcontainingCOMMAND_RESULT:,STATE_DELTA:,ANCHOR: Active Spec is '<path>'. Step <N>., andNEXT:lines. Total payload capped under ~500 chars for typical commands.
The hooks never read or write the active spec body itself — only the pointer and the scratchpad. This is what makes them safe to run on every session and every Bash call.
Phase: First-run welcome read hydrate.signal is_first_run field; on true:
render welcome block + choice 1/2/3;
on false: skip silently
Phase: Fingerprint (silent, internal) → state/local-symbols.json + template surfacing
Phase: Wizard 4 mandatory §8.2 questions; cached by author-spec-hash
Phase: Intent (user input)
Phase: Feasibility refuse if physically impossible
Phase: Walker loop init-or-resume; peek-pending → question → answer-concern
→ yield-check; loop until stop
Phase: Draft render from state.answered → .spec.md.draft
atomic write; Confirm: yes / refine / cancel
Phase: Evaluator gate — setup wizard ensure ~/.spectre/reviewer.toml exists
Phase: Evaluator gate — spec evaluation ── see "Evaluator pipeline" below ──
Tier 1 spec_ast (AST)
Tier 2 coverage_gate (structural)
Tier 3 llm_judge (DeepSeek, opt-in)
Halt on any block-severity finding.
Phase: Evaluator gate — ADR generation (conditional)
decisions/<NNNN>-<slug>.md per Decision marker
Phase: Lock atomic rename .draft → .spec.md
flip specs/.active
reset state/scratchpad.json (v2 shape)
write <slug>.spec.md.eval.json sidecar
clear state/.eval-bundle.json
write <slug>.envelope.json handoff envelope
Phase: Transition print VISION LOCKED signal
Two architectural commitments shape the protocol:
- Draft → confirm → lock is explicit. No silent locks. The user reads the draft on disk in their own editor before saying yes.
- The evaluator's bundle is materialized once. §6.4 builds a
ReviewBundlecontaining preview ADRs, preview Resources, and tier classifications, persists it tostate/.eval-bundle.jsonkeyed by the draft's SHA-256, and §6.5/§6.6/§6.7 read from it. No recomputation; if the draft changes between steps, the SHA mismatch forces a re-run.
Phase: Mode routing parse check / auto / track args
Phase: Track default = "default"; create track if absent
Phase: Tier 0 envelope validate <slug>.envelope.json before reading spec
Halt: status=tampered / schema violation
Phase: Context read .active + scratchpad[track]
Halt: SPEC COMPLETE if all steps verified
Phase: Environment ensure_venv + normalize_action rewrites
Halt: VENV CREATION FAILED (no fallback)
Phase: Pre-flight re-run prior step's verification
Halt: ROOT-STATE DESYNC on fail
Phase: Check mode verify-only, no scratchpad write
Phase: Tier classifier spectre tier — silent/repo/host/network/NA
Halt+ask on host, network, or never-autonomous
Post-halt-success: adopt / once-only / never-ask-again
Sandbox-paradox brake at 3 adoptions/session
Phase: Resource acquire spectre track acquire via UDS to supervisor
Halt: RESOURCE QUEUED if at capacity
Phase: Reasoning emit print "WHY: <why text>"
Phase: Execute Bash → PostToolUse hook fires compact.py
Phase: Verify run verification:; halt on non-zero
State Auditor: informational PBT-lite checks
Phase: Branch on verification
Path A (pass): advance step, CDLC ledger implement entry
Path B (fail): one Option-B retry with diagnosis, then halt
Phase: Drift every 5 successful steps; re-read §1, audit next batch
Halt: DRIFT DETECTED on concern
Phase: Resource release spectre track release on terminal step state
Phase: Failure log append to scratchpad.failed_hypotheses[]
Phase: Finding capture Path B retry succeeded → prompt project/spectre
The persistence-tier gate (§3.5) is the single source of truth for halt-vs-execute. v1 used a regex risk-gate inline in the skill prose; v0.2.1 replaced it with bin/tier.py so the rule set is testable and project-overridable.
v0.3.1 NEVER_AUTONOMOUS additions (closes #1 gap 6): systemctl <verb> (start/stop/restart/reload/enable/disable/mask/unmask, with --user/--system flag tolerance), loginctl enable-linger / disable-linger, hostnamectl set-*, timedatectl set-*, sysctl -w. v0.3.0 missed these because the verb regex only checked the binary name; the v1.1.0 BTC proxy test run hit three host-state-mutating systemctl invocations that the classifier treated as silent until the agent applied judgment-override.
v0.3.1 loopback downgrade (closes #1 gap 9): _is_network() now parses URLs in argv. 127.0.0.1, localhost, [::1], 0.0.0.0, and RFC1918 (10.*, 172.16-31.*, 192.168.*) downgrade curl/wget to path-tier classification on the output file. The packet never leaves the kernel; halting is pure friction. Variable URLs ($VAR) keep network tier — false-positive is the safe default.
draft.spec.md.draft
│
▼
build_bundle() ← spec_evaluator.py
│
├──→ Tier 1: spec_ast.classify(draft_path)
│ Pure parse/structure/tautology checks.
│ Imports NOTHING from bin/tier or bin/resources.
│ Findings: missing-why, soft-verification,
│ action-not-probed, missing-receiver-calibration.
│ v1.0 addition: _v1_structural_checks() verifies §§8.3-8.7
│ + §§9-13 present (or N/A), validates contract subsections,
│ warns on more-than-two N/A views. Gated by is_v1_spec().
│ Short-circuits for pre-v1.0 specs (no Spec-version: frontmatter).
│
├──→ Tier 2a: coverage_gate.classify(draft_path,
│ preview_adrs=preview_adrs)
│ Cross-checks §8.1 calibration vs action path captures.
│ Findings: undeclared-resource (warn),
│ undeclared-host-path (block),
│ calibration-hard-violation (block),
│ decision-without-adr (warn, deterministic).
│
├──→ Tier 2b: cross_view_gate.classify(draft_path) [v1.0 only]
│ Cross-view consistency checks; short-circuits for non-v1.0 specs.
│ 1. Resolves §§9-13 cross-view string references against §8.x fields.
│ 2. Validates exemplar bindings against the metis catalog.
│ 3. Enforces taxonomy-version match (exemplar vs catalog).
│ 4. Flags fingerprint↔hard-contract contradictions across §§8.3-8.7.
│ Entry point: cross_view_gate.classify() — never raises.
│
└──→ Tier 3: llm_judge.evaluate(spec_text, config=…)
Single JSON-only API call to DeepSeek deepseek-v4-flash.
Contradiction-tuple protocol — 10 kinds + unrecognized fallback.
v1.0 addition: _build_exemplar_context() — when the spec binds
exemplars in §§9-13, conventions lists are injected into the
contradiction prompt (capped at 5 violation tuples per call).
No additional API calls — context appended to the single call.
CoT faithfulness cite-and-verify pass (zero extra API calls
when no block tuples exist).
Budget instrumentation: emits one _status.emit("info", "tier3.budget", ...)
per call: `INFO tier3.budget calls=1 exemplars_injected=N dismissals_by_fp={…}`
(v1.1: suppressible via SPECTRE_QUIET=1).
Never raises. All failures become a single
tier3-unavailable info-severity sentinel finding.
│
▼
ReviewBundle persisted at state/.eval-bundle.json
│ keyed by draft SHA-256
▼
evaluate(draft, config_path, persist_dir) → EvaluatorResult
│ - apply severity overrides (one-way: raise only)
│ - filter dismissals by stable fingerprint
│ - return findings + sidecar_payload
▼
/vision §6.4 halts on block-severity findings; otherwise continues to §6.5.
Why three tiers?
- Tier 1 is deterministic and free. Catches the cheapest mistakes (missing
why:, soft verifications) before any LLM call. v1.0 extends this with structural checks for §§8.3-8.7 and §§9-13. - Tier 2 is structural and free. v1.0 splits this into two modules: (a)
coverage_gatecross-checks the §8.1 hard contract against actions' path captures — this is where most real bugs hide; (b)cross_view_gatecross-validates the six-view model: reference resolution, exemplar catalog integrity, taxonomy-version consistency, and fingerprint↔hard-contract coherence. - Tier 3 is semantic and paid. DeepSeek is structurally other-than-Claude — different training distribution, different blind spots, different priors. The adversarial-reviewer principle: same-family review has correlated blind spots; different-family review breaks the correlation. Costs are capped (
budget_tokens_per_spec) and opt-in (enabled = falseby default).
Stable fingerprints, dismissable findings. Every finding has a SHA-256 fingerprint over {tier, kind, scope, step, steps, ref} — message text is deliberately excluded so LLM nondeterminism doesn't break dismissals. A user dismisses a Tier 3 finding by appending # tier3-dismissed: <fingerprint> "<reason>" to the spec; the next evaluator run skips that finding. Tier 1 and Tier 2 findings are dismissable=False — structural rules don't get to be argued away.
bin/walker.py is the strict-hybrid state machine that drives /vision Steps 1–5. It owns walk state, branching, dependency-tracked invalidation, and stop conditions. The skill phrases the walker's structured Concerns into natural-language questions; the walker is canonical, the rendered question is best-effort.
Stop conditions (any triggers halt; evaluated in this deterministic order):
- Author types
stop→stop_reason = "author-arbitrated". - Tier 3 yield-delta converged (3 consecutive rounds <2 new findings).
- Max-rounds hit (default 30, configurable in
~/.spectre/walker.toml). - Per-receiver exhaustion (no non-stale pending concerns remain).
State persists at state/.walk.json (atomic JSON write — same pattern as bin/_scratchpad.atomic_write). A walk in progress survives session interruption. Revising an earlier answer marks all transitively-dependent concerns stale; the walker skips stale concerns in next_concern. The author sees the invalidated set as a diff and chooses re-walk or accept-stale.
init_walk seeds five concerns: one assumption-surface (round 1: surface unstated assumptions in the intent) plus four receiver-clarification concerns covering the §8.1 hard contract (mutates, never-touches, decision-budget, reboot-survival). The downstream evaluator (§6.4) requires these fields; seeding them at walk-init guarantees they're answered before draft materialization.
v1.0 WalkState additions. The WalkState dataclass gains:
view_scope: dict[str, str]— tracks in-scope / not-applicable per non-agent view. Keys:"product-input","product-output","human-user","integrator","operator". Values:"in-scope"or"not-applicable". Set by answering the five scope-check concerns (scope-product-input, etc.) mapped in_VIEW_SCOPE_CONCERN_IDS.product_input_asked: bool,product_output_asked: bool,human_user_asked: bool,integrator_asked: bool,operator_asked: bool— idempotency guards for the per-view concern generators.
v1.0 generator inventory. Five new generate_*_concerns() functions (one per non-agent view) extend the existing generator set:
| Generator | Scope concern ID | Follow-up concern count |
|---|---|---|
generate_product_input_concerns() |
scope-product-input |
4 |
generate_product_output_concerns() |
scope-product-output |
4 |
generate_human_user_concerns() |
scope-human-user |
3 |
generate_integrator_concerns() |
scope-integrator |
4 |
generate_operator_concerns() |
scope-operator |
4 |
Each generator emits the scope-check concern first. N/A scope short-circuits follow-ups; in-scope answers cascade into the view's concern set.
WALKER_VERSION = "1.0.0" — hard cutover. Pre-v1.0 state files (walker_version != "1.0.0") raise ValueError on load with a recovery hint. Operators must delete state/.walk.json and re-run /vision. No migration path; no version dispatch. This is intentional: the new view_scope and five per-view *_asked fields cannot be defaulted from an older state shape without silently skipping view concerns for in-flight specs.
Reference:
- Design:
docs/superpowers/specs/2026-05-06-spectre-v0.4-cdlc-closure.md - Plan:
docs/superpowers/plans/2026-05-06-v0.4.0-walker.md
bin/substrate_wizard.py fires at /vision Step 0.5 to populate §8.2 via 4 mandatory questions. v1.0 extends this with run_per_view(), which renders one §8.x block from validated flag values.
Per-view fingerprint vocabularies (_VIEW_FINGERPRINTS). Each of the six views has its own allowed receiver-fingerprint set:
| View | Vocabulary |
|---|---|
implementing-agent |
claude-code+human, claude-code-autonomous, non-claude-ai, human-only |
product-input |
human-typed, programmatic-trusted, programmatic-untrusted, streamed-event, not-applicable |
product-output |
human-reader, programmatic-consumer, streaming-sink, log-aggregator, not-applicable |
human-user |
cli-power-user, cli-novice, gui-only, no-human-user, not-applicable |
integrator |
library-consumer, api-consumer, webhook-subscriber, sdk-author, no-integrator, not-applicable |
operator |
on-call-engineer, sre-team, self-operated, no-operator, not-applicable |
Per-view trust-profile vocabularies (_VIEW_TRUST_TOKENS). Trust tokens are scoped to each view and isolated by design. A token valid in one view's vocabulary is rejected if specified in another view's profile — cross-vocabulary cross-pollination would mislead the evaluator. Example: handles-secrets is valid for implementing-agent; accessibility-required is valid for human-user. Neither is accepted in the other's profile field.
run_per_view() entry point. Caller (/vision skill) walks the view_scope dict from WalkState and invokes run_per_view(view=..., receiver=..., trust_profile=..., binding=...) once per in-scope view. Returns a rendered §8.x markdown block. When receiver == "not-applicable", the block degenerates to a single not-applicable: <reason> field. Cache responsibility is on the caller (see run_with_flags for the §8.2 cache pattern).
When a v1.0 spec binds exemplars in §§9-13, _build_exemplar_context() appends the bound conventions lists to the Tier-3 contradiction prompt:
- Walks §§9-13 sections and extracts
exemplar:<view-type>:<slug>references. - Looks each reference up in the metis catalog via
_catalog.lookup(). Missing entries are silently skipped (Tier-2 cross-view gate already flagsexemplar-not-found). - Formats the conventions list per binding and appends to the user-message portion of the API call. Capped at 5 violation tuples per call — same single API call, no multiplied requests.
- DeepSeek is instructed to identify steps whose output would violate a bound exemplar's conventions.
Budget instrumentation. Every evaluate() call emits one _status.emit("info", "tier3.budget", ...) line, regardless of whether exemplars were injected:
INFO tier3.budget calls=1 exemplars_injected=N dismissals_by_fp={...}
calls is always 1. The ship-gate harness uses this line to confirm Tier-3 call volume stays within budget. v1.1 format change: emission moved from raw print(... json.dumps(...)) to _status.emit("info", ...), gaining SPECTRE_QUIET=1 suppression for free and consistency with every other reviewer status line. Parsers should split on whitespace + key=value rather than json.loads the tail.
Pre-1.0 state files (walker_version != "1.0.0" in state/.walk.json) raise ValueError on walker.load() with a message of the form:
walker_version mismatch: file has '0.4.1', walker is '1.0.0';
remove state/.walk.json and re-run /vision to start a fresh walk.
No migration tool is provided. The state-shape change (five new per-view fields + view_scope dict) cannot be safely defaulted from a v0.x state file without silently omitting view-concern families for any in-flight spec that would have answered them. Recovery: delete state/.walk.json and re-run /vision.
bin/observations.py and bin/personal_rules.py close two of the three remaining CDLC legs (Distribute is v0.4.2).
Observe. Every TIER GATE halt in /implement records a structured row to ~/.spectre/observations.jsonl: timestamp, fingerprint, classifier_label, project_path, spec_slug, action. The log is append-only and per-user (not per-project) so recurring halt patterns surface across all the user's Spectre projects. find_recurrences(threshold=N) returns fingerprints recurring ≥N times — consumed by v0.4.2's Adapt template-patch flow.
Adapt. ~/.spectre/personal-rules.toml is the per-user TOML override store. The /implement post-halt-success prompt (§3.5b) is the only sanctioned writer. bin/tier.should_halt() consults personal_rules.is_classifier_halt_overridden() for every host/network-tier halt:
never_autonomous_matchis non-overridable — those rules are never downgraded.- If the action's classifier reasons reference any path in the active spec's §8.1 hard contract (
spec_locked_paths, parsed viabin/coverage_gate.parse_81_block), the personal-rule cannot override — spec rules are immune. - Otherwise, if the
(classifier_label, fingerprint)pair has an entry in personal-rules.toml, the halt is downgraded.
Sandbox-paradox brake. Per research/developing-a-safe.md (vault), HITL approval gates create permission-fatigue (users rage-bypass safety). v0.4.1 caps adoptions at 3 per session before the post-halt prompt stops firing, requiring the user to manually review personal-rules.toml. The counter is persisted per-track in state/scratchpad.json["tracks"][<track>]["session_adoption_count"] — surviving the per-heredoc Python subprocess fork that would otherwise reset module-state. Reset only by manually clearing the field or via the test helper personal_rules.reset_session_adoption_count_persistent().
Reference:
- Design:
docs/superpowers/specs/2026-05-06-spectre-v0.4-cdlc-closure.md§6.3, §6.4 - Plan:
docs/superpowers/plans/2026-05-06-v0.4.1-observe-adapt.md
bin/cdlc_ledger.py, bin/templates.py, and bin/template_patcher.py close the third leg of the CDLC.
Ledger. Every Generate→Test→Lock→Implement→Halt→Adapt transition is appended to per-project state/cdlc-ledger.json via atomic write. Read-only audit surface — no user-facing command. Call sites: /vision §6.7 (lock=generate), /implement §6 Path A (implement) and §3.5 (halt), bin/observations.record_halt (halt), bin/personal_rules.append_adoption (adapt).
Distribute. ~/.spectre/templates/{specs,skills}/ is the per-user template store. bin/templates.import_template copies a stored template into a new project (specs land at ./specs/<name>.spec.md.draft so the /vision flow still gates the lock; skills land at ./skills/<name>.md). bin/templates.export_template is the reverse. Local-only — remote sync is still deferred (not in v0.5 or v0.6; tentatively v0.7+).
Adapt-patches. When observations.find_recurrences(threshold=3) returns recurring halt fingerprints AND those fingerprints aren't already covered by personal_rules, bin/template_patcher.detect_patch_candidates lists them and template_patcher.propose_patch writes a markdown patch to ~/.spectre/template-patches/proposed/<slug>.md. SessionStart's bin/hydrate.surface_pending_template_patches reports the count; bin/hydrate.detect_and_propose_patches writes new proposals at session start. Manual accept/reject only.
Deferred-prompt durability. bin/_scratchpad.track_default() gains pending_adoption_prompt: dict | None. /implement §3.5 writes this on TIER GATE halt + user=yes; §3.5b reads it post-Path-A and clears FIRST (before adopt-write) so an adopt-write failure does not strand the prompt for replay next session.
Reference:
- Design:
docs/superpowers/specs/2026-05-06-spectre-v0.4-cdlc-closure.md§6.5, §6.6 - Plan:
docs/superpowers/plans/2026-05-06-v0.4.2-cdlc-distribute.md
For multi-track projects (/implement payments, /implement notifications), Spectre runs a per-project Unix domain socket daemon at runtime/supervisor.sock:
- Spawn-on-demand. First
/implement <track>call detects no live supervisor, double-forks one, waits for the socket, then dials in. - Idle self-shutdown. No requests for 30 minutes → daemon exits.
- Reboot recovery. PID-file +
/proc/<pid>/statactor fingerprinting reaps stale locks left by SIGKILL'd or rebooted-out actors. Reconcile is idempotent. - Single-threaded
select()loop. No concurrency primitives, no thread pools, no race conditions to debug. Stdlib only.
Resource nodes live in specs/.graph.md:
---
id: res-port-8080
type: resource
title: HTTP server bound to TCP 8080
status: active
edges: []
---Steps reference them via resources: [res-port-8080]. The supervisor grants the lock (or queues the requester) before §3.7 fires the WHY emit and the action runs. Released after §6 verification passes — or on terminal halt.
Pre-v1.0 shape (§§1-8.2):
§1 Hard Problem one-paragraph non-obvious challenge
§2 First Principles 3–7 bullets (physics/logic, no analogies)
§3 Algorithm Audit Delete / Simplify / Accelerate
§4 Speed-of-Light Limit one paragraph: the physical ceiling
§5 Physics Guardrails system invariants
§6 Steps YAML list: why/action/verification (+ properties, resources)
§7 Success Criteria binary checklist
§8 Receiver Calibration 8.1 hard contract (machine-enforced)
8.2 cognitive-substrate contract (§8.2 wizard-prompted)
v1.0 additions (§§8.3-8.7 + §§9-13):
§8.3 Product-input substrate per-view receiver-fingerprint + trust-profile + contextual-binding
§8.4 Product-output substrate (same schema; separate vocabulary)
§8.5 Human-user substrate (same schema; separate vocabulary)
§8.6 Integrator substrate (same schema; separate vocabulary)
§8.7 Operator substrate (same schema; separate vocabulary)
§9 Product-Input View contracts: mechanical | coverage | exemplar-bindings
§10 Product-Output View (same three-type taxonomy)
§11 Human-User View (same three-type taxonomy)
§12 Integrator View (same three-type taxonomy)
§13 Operator View (same three-type taxonomy)
v1.0 specs must open with **Spec-version:** 1.0 frontmatter. Any view may be marked not-applicable: <reason> to degenerate to a single-field block; more than two N/A views generates a warn-severity excessive-not-applicable finding.
§8.1 is the load-bearing addition from v0.3.0. The four required fields — mutates, never-touches, decision-budget, reboot-survival — are cross-checked by Tier 2 against every action's path captures. This is what catches "the spec said it would only touch /opt/foo/ but step 7 writes to /etc/" before any code runs.
| Failure | Mitigation | Test |
|---|---|---|
| Broad PostToolUse matcher fires on every tool | matcher: "Bash" exact; not regex |
test_compact.py |
| Hydrator dumps full scrollback into context | hydrator emits only spec body + STATE line | test_hydrate.py |
| Verification fail loops forever | one Option-B retry, then hard halt | test_e2e.py |
| Concurrent writes to scratchpad corrupt JSON | atomic write via mkstemp + os.replace | test_scratchpad.py |
additionalContext payload bloats |
length-cap + path FIFO at 200 entries | test_compact.py |
| Stale lock from SIGKILL'd actor blocks track | /proc/<pid>/stat fingerprint reconcile |
test_supervisor.py |
| Drift across 20+ step specs | every-5-steps drift checkpoint re-reads §1 | test_e2e.py |
| LLM finding dismissals broken by nondeterminism | fingerprint excludes message text | test_findings.py, test_dismiss_integration.py |
| Severity downgrade in user config | validate_no_severity_downgrade raises ValueError |
test_eval_metadata.py |
| Tier 3 outage crashes evaluator | all errors → tier3-unavailable info sentinel | test_llm_judge.py |
Three principles drive every choice:
- Determinism first. Anything that can be a deterministic check (Tier 1, Tier 2, the persistence-tier classifier, atomic file ops) is one. LLM calls happen at one well-defined gate, not scattered through the protocol.
- Atomic transitions. Every state change is
mkstemp + os.replaceor equivalent. There is no torn-write window where the active spec is half-flipped or the scratchpad is half-updated. - Adversarial review at the spec layer. Code review catches code bugs; spec review catches spec bugs. v0.3.0 ships the spec-review layer because the cost of a wrong spec is N steps of wrong implementation, and the cheapest place to halt is before lock.
For the historical context behind specific decisions, see docs/superpowers/specs/ (architecture briefs) and docs/superpowers/plans/ (implementation plans). Both directories are archival — they record what was built and why, not what to do next.
Hard problem: skill prose contained 20 python3 - <<'PY' ... PY heredoc blocks that forked fresh interpreter processes, bypassed argparse schemas, and let slug substitutions, path drift, and inline Path(...) constructions silently diverge from the underlying bin/ functions they called. Any bug in the heredoc was invisible to the test suite.
Design decision: Every heredoc becomes a python3 -m bin.<module> <subcommand> invocation against a versioned CLI surface. No new business logic — only __main__ entry points wrapping existing public functions. The CLIs ship in PRs #18 (Phase 2A), #19 (Phase 2B), #20 (Phase 2C), and #33's predecessor (Phase 2D).
Load-bearing files:
bin/cdlc_ledger.py,bin/observations.py,bin/_scratchpad.py,bin/personal_rules.py,bin/track.py,bin/adr.py,bin/templates.py,bin/setup_wizard.py,bin/walker.py— all gained__main__entry points in v0.5.0.tests/test_skill_prose_no_heredoc_python.py— drift-prevention guard; per-file ceilings tightened to zero. Any futurepython3 - <<'PY'inskills/**/SKILL.mdbreaks CI immediately.
Cumulative effect: 20 heredocs gone, ~192 LOC removed from skill prose, 928 tests at release.
References: Issue #13 (closed), docs/superpowers/audits/2026-05-06-issue-13-heredoc-audit.md.
Hard problem: the v0.5.1 live test of yt-readable surfaced five gap classes (A–E) where the pre-lock evaluator returned PASS but /implement auto halted on bugs the evaluator should have caught: uncreated artifacts (Gap A), import-before-install (Gap B), scaffold-without-implementation (Gap C), unparseable Python in verifications (Gap D), and PEP 668 / venv isolation (Gap E). Per Copilot/GPT-5.4 peer review (#32): the fix is deterministic contracts + executor-owned environment + hard gating, not prose-inferred graphs.
Design decisions:
-
Executor-owned venv (
bin/managed_venv.py). The implementor creates and ownsstate/.venv/(mode 0700) rather than relying on the system Python.normalize_actionrewrites action head-tokens to use the venv Python, preserving shell operators byte-identical. Stalepyvenv.cfg(e.g. after Python upgrade) triggers HALT rather than silent misbehavior. -
Explicit step contracts (
produces:/requires:). Eight contract types:file:,package:,console-script:,route:,module:,binary:,db-table:,db-column:. Tier 1 cross-validatesrequires:against priorproduces:. Mismatch → blockunowned-requirement. Steps with no contracts → warnmissing-contract(backward-compatible).contract_resolutionblock added to.eval.jsonsidecar. -
Tier 1 deterministic gap-closers.
verification-syntax-error(block) — everypython3 -c "<body>"compile-checked at lock time.action-invokes-uncreated-artifact(block) — absolute paths undermutates:with no prior authoring step.unowned-requirement-heuristic(block) — curl routes, HTML tags, SQL columns, Python imports with no prior owner. Allowlists for universal probes andproduces:declarations shadow heuristics. -
Tier 3 contradiction-tuple protocol. DeepSeek system prompt rewritten (~540 tokens) to force JSON-only output. Ten contradiction kinds +
unrecognizedfallback. Single API call replaces the prior three-prompt prose loop.DEEPSEEK_MODELdefault changed fromdeepseek-reasoner→deepseek-v4-flash— structured I/O protocol doesn't need reasoner-style prose output.
Load-bearing files: bin/managed_venv.py (new), bin/spec_ast.py (contracts + gap-closers), bin/llm_judge.py (tuple protocol), bin/eval_metadata.py (contract_resolution in sidecar), bin/spec_evaluator.py (DEEPSEEK_MODEL default).
References: Issues #31 (gap classes), #32 (design brief + Copilot/GPT-5.4 review), PR #33.
Hard problem: five vault concept pages mapped onto Spectre's weak points after v0.5.2 identified that the bytewise integrity of the vision→implement handoff was not enforced: the sidecar could be tampered or the spec body replaced without the executor noticing (Gap E from the v0.5.2 essay-followup).
Design decisions:
-
Handoff envelope (
bin/handoff_envelope.py+bin/handoff_validator.py). A JSON-Schema-validated envelope wraps the vision→implement handoff. Schema:protocol_version,receiver,spec_path,sidecar_path,policy_hash,spec_sha256,sidecar_sha256,contract_resolution,walker_yield_history,walker_stop_reason,decisions_indexed,integrity_hash,created_at. The integrity hash covers the actual artifact bytes (spec.md + sidecar.eval.json), not just envelope metadata.Step 0.7 in
/implement(skills/implement/SKILL.md) is the new Tier 0 check, inserted between Step 0.5 (track selection) and Step 1 (read context). Four outcomes:envelope-missing(warn — pre-v0.6 spec, allow),envelope-tampered(block — content modified after lock),envelope-malformed(block — schema violation), clean (proceed). -
Walker yield countdown.
bin/walker.pyemits prediction-ready status lines:"YIELD: round N added M new T3 findings; stopping when last K rounds all <T (currently: [a,b,c])"instead of raw delta numbers. Newnegative-pathconcern kind +generate_negative_path_concernswith idempotency guard. -
Negative-path Tier 1 enforcement (
bin/spec_ast.py). New optionalnegative-paths:block per step (list of{trigger, handler}dicts). Tier 1 warnsmissing-negative-pathwhenproduces:is non-empty andnegative-paths:is absent. Blocks whenreboot-survival: required(data-loss hazard). Malformed-only entries underreboot-survival: requiredalso escalate to block. -
Tier 3 CoT faithfulness (
bin/llm_judge.py). A single batched cite-and-verify pass runs after the primary contradiction tuples. Block-severity tuples (missing-producer,shallow-ownership) are demoted to warntier3-unfaithful-contradictionif DeepSeek can't cite supporting spec text (case-insensitive substring). Parse failure → conservative: keep block, appendtier3-faithfulness-malformedwarn. Zero extra API calls when no block tuples exist.
Load-bearing files: bin/handoff_envelope.py (new), bin/handoff_validator.py (new), bin/eval_metadata.py (write_envelope_alongside_sidecar()), bin/walker.py (yield countdown + negative-path concerns), bin/spec_ast.py (negative-paths enforcement), bin/llm_judge.py (faithfulness pass), skills/vision/SKILL.md (Step 6.7 envelope write), skills/implement/SKILL.md (Step 0.7 Tier 0 check).
References: vault pages concepts/context-as-cognitive-substrate.md, entities/standardized-handoff-envelope.md, entities/context-sled.md, entities/handoff-validator.md, entities/planner-generator-evaluator-triad.md, research/cot-monitorability.md.
Hard problem: skill prose's python3 -m bin.X invocations are run from the user's project cwd, where bin/ may not be on sys.path. Without an explicit PYTHONPATH="${CLAUDE_PLUGIN_ROOT}" prefix, plugin-internal module resolution silently fails at runtime in any project that doesn't happen to have a bin/ directory at cwd (issue #30).
Fix: every python3 -m bin.X invocation in skills/vision/SKILL.md and skills/implement/SKILL.md now carries the PYTHONPATH="${CLAUDE_PLUGIN_ROOT}" prefix. A PYTHONPATH note section at the top of both skill files explains the requirement. CI sentinel tests/test_skill_pythonpath_consistency.py scans all skills/**/SKILL.md bash code blocks and asserts every python3 -m bin.X line carries the prefix.
References: Issue #30 (closed).
Hard problem: the v0.6.1 retest of an activity-ingestion-daemon spec surfaced two bugs that defeated whole evaluator subsystems silently. (1) Tier 1's heuristic shadow couldn't suppress an unowned-requirement-heuristic block when a step verified from spectre_daemon.blocklist import is_blocked because _PYTHON_IMPORT_ALT_RE was matching the SYMBOL is_blocked as a module name — unmatched against declared_modules, fired anyway. (2) Tier 3 silently degraded for any user whose ~/.spectre/reviewer.toml predated v0.5.1: stale model = "deepseek-reasoner" returned HTTP 401 against the user's plan, and llm_judge reported it as socket-timeout — DeepSeek unreachable, sending debug effort toward network instead of credentials.
Fixes (issues #36, #37):
- Contract-shadow precision.
bin/spec_ast.pynow skips_PYTHON_IMPORT_ALT_REmatches whose start lies inside a span already matched by_PYTHON_IMPORT_RE(thefrom X import Yform). Parent-prefix match is added to both contract resolution (unowned-requirementblock check on declaredrequires:entries) and the heuristic shadow:package:foosatisfiesmodule:foo.bar;module:foo.barsatisfiesmodule:foo.bar.baz. Three regression tests intest_spec_ast_v052_gaps.py. - Stale reviewer.toml auto-migration.
bin/setup_wizard.maybe_provision()no longer treats every existing config as authoritative. It detects (a)model in {deepseek-reasoner, deepseek-chat}, (b) missingchunk_timeout_s, (c) missingtotal_timeout_s. Any one trips migration: backup written toreviewer.toml.bak-<timestamp>(mode 0600), config rewritten with current defaults (deepseek-v4-flash+ chunk/total split timeouts), preserving the user'senabledflag andapi_key_env. Returns new outcome"migrated". Stderr breadcrumb points at the backup. Five regression tests intest_setup_wizard.py. - Tier 3 error-class disambiguation.
bin/llm_judge.evaluate()now classifiesurllib.error.HTTPErrorexceptions:401/403→"auth failure (HTTP NNN — check ~/.spectre/secrets.env or DEEPSEEK_API_KEY)",400→"bad request (model X may be unavailable on your plan)",5xx→"provider error (HTTP NNN)", anything else →"http-NNN". Network/timeout paths unchanged. Four regression tests intest_llm_judge.py. - Auth-failure prominence.
skills/vision/SKILL.mdStep 6.4 now requires a⚠ Tier 3 unavailable due to auth — fix ~/.spectre/secrets.env or DEEPSEEK_API_KEY then re-run /vision.banner ABOVE thetier 1/2/3status block whenever thetier3-unavailablefinding's message contains the substringauth failure. The banner makes credential issues actionable without scanning the findings list.
EVALUATOR_VERSION = "0.6.2". 1280 tests passed at release.
References: Issues #36, #37 (closed). PR #39, release v0.6.2.