A Claude Code marketplace with one plugin, dev-workflow, that installs a spec-driven
workflow for building software with agents. It adds three mechanisms: two independent
cross-model review gates, a self-hardening findings ledger, and repo-enforced quality.
- Two independent review gates. A different model than the author reviews the design before it becomes code (Gate A) and the diff before it lands (Gate B).
- A self-hardening ledger. Every finding is fingerprinted, so recurrence of a class is detectable — and a recurrence escalates the response one rung harder: prose note → lint rule → type constraint → test.
- Repo-enforced quality. One quality command runs in CI on every change, so nothing red can merge.
Why each of these, and how to adapt them: docs/coding-workflow.md.
dev-workflow:intake |
skill — a raw idea or voice transcript (German or English) becomes a reviewable story. Captures WHAT and WHY; refuses to invent the parts that aren't there. |
dev-workflow:harden-finding |
skill — one review finding becomes a lint rule, type constraint, test, or documented convention, at the right rung, recorded in the ledger. |
/dev-workflow:process-pr-review |
command — validates PR bot comments against the code and your invariants, replies to each, fixes regressions, tracks pre-existing issues. |
dev-workflow:finding-triage |
agent — read-only, fresh context, judges whether one PR-bot claim is actually true of the code. Used by the PR processor; never counts as a review gate. |
/dev-workflow:workflow-init |
command — scaffolds the per-project files, then interviews you to write AGENTS.md. |
| codex-gate hook | non-blocking reminders that count Gate A and Gate B passes, and verify a Gate-B review against a fingerprint of the effective index plus the included worktree content, as of the hook's invocation — a deliberate superset of any one commit's payload, so the gate errs toward firing. Always exits 0. |
1. Install superpowers, then this plugin:
claude plugin marketplace add obra/superpowers-marketplace
claude plugin install superpowers@superpowers-marketplace
claude plugin marketplace add dsnger/dev-workflow-kit
claude plugin install dev-workflow@dev-workflow-kitTo update later:
claude plugin marketplace update dev-workflow-kit
claude plugin update dev-workflow@dev-workflow-kit # qualified form is a workaround: the bare
# name is documented but errors "Plugin
# 'dev-workflow' not found" (CLI 2.1.x)Updates arrive only when the plugin version is bumped, and a running session keeps the
old version until you restart it or run /reload-plugins.
2. Prerequisites:
-
superpowers — not vendored. The workflow's middle is its skills; without it,
intakehands off to nothing and Gate A never fires. -
Codex — the reviewer behind both gates. Needs the Codex CLI, authenticated with an OpenAI account — a real external dependency, not just the
.mcp.jsonentry/workflow-initwrites for you. It also needs a Codex MCP server that exposesexecandreview— the gates and their pass counters key on those two tool names. Use themcp-codex-devserver/workflow-initpins, which has both. The officialcodex mcp-serveris a different server exposing a singlecodextool, which can't be attributed to Gate A (reviews text) or Gate B (reviews a diff): with it connected, no pass ever counts and Gate B reports "not run" forever (the hook says so, once).Both gates have Codex write its findings to a file under
.context/codex-reviews/and reply with a one-line summary, because a long finding list returned inline can come back cut off — and a cut between findings looks exactly like a short list. Claude Code limits MCP tool output to 25,000 tokens by default;MAX_MCP_OUTPUT_TOKENSraises that, which is worth trying as a secondary mitigation but is not the fix. It moves the ceiling rather than removing it, and it applies only to tools that don't declare their own text limit viaanthropic/maxResultSizeChars— whethermcp-codex-devdeclares one is not established here. The file protocol is what actually keeps findings out of the response. -
gh— optional; only/dev-workflow:process-pr-reviewuses it.
3. Run /dev-workflow:workflow-init in each project. It verifies the rest and tells
you what's missing — git repo, superpowers, Codex (not configured / not loaded / ok),
gh, AGENTS.md, stack — before writing a single file. Then follow
docs/getting-started.md for your first story.
idea → intake → brainstorm → spec → Gate A → plan → Gate A → implement → quality
battery → Gate B → PR → process-pr-review → merge, running every finding worth
keeping through harden-finding.
New to the workflow? docs/getting-started.md walks one
feature through every step — what you do, what happens, what you see.
The hook only speaks up in initialized projects — the ones where /workflow-init
has run (it writes .context/codex-gate.on, and its CLAUDE.md §5 counts too). The
plugin is installed once per machine; every other repo you open hears nothing from it.
Per-workspace knobs, all files under .context/:
| codex-gate.floor | a positive integer; moves the 3-passes-per-gate floor. |
| codex-gate.off | silences the reminders; state keeps tracking, so re-enabling is accurate. |
| codex-gate.tools | execTool=<name> and/or reviewTool=<name> — counts a Codex server whose tools aren't named exec/review, and only worth it if that server really does separate text-review from diff-review; aiming both gates at one general-purpose tool moves the counters while neither gate means what it says. Unparseable lines are ignored, so a typo can't quietly unhook a gate. |
Without Codex, /workflow-init degrades honestly instead of scaffolding gates that
can't run: it silences the hook and marks CLAUDE.md §5 INACTIVE with the re-enable
path. There is deliberately no same-model fallback reviewer — explicitly gateless beats
implicitly self-reviewed
(why).
Five things stay yours on purpose — a plausible default nobody verified is worse than an honest gap (reasoning):
- Quality-battery tools — the plugin names the roles (typecheck, lint, dead code, duplication, tests), never the tools. You wire them into one command.
- Baselines — generated from your repo, in their own commit.
- Custom lint rules — two ship as examples only
(
examples/eslint-rules/); they encode one stack's invariants and won't transfer. - Branch protection — a repo setting, not a file; worth doing once the battery runs.
AGENTS.md— your invariants. Both gates check against it, so a generic one makes "check against our invariants" read as satisfied when nothing was.
CI (.github/workflows/ci.yml) runs four checks on every
PR and push to main: shellcheck --shell=sh over all three executables and their test
files, the hook's test suite,
scripts/check-invariants.sh (invariants 5 and 6) plus
both checkers' regression suites, and claude plugin validate . --strict.
A fifth check runs on pull requests only:
scripts/check-version-bump.sh (invariant 12), which
needs a base branch to diff against. Its suite runs unconditionally with the others;
the checker itself does not. That is sufficient here because main accepts changes only
through a PR: a branch ruleset (protect-main., active since 2026-07-19) requires a
pull request and a green quality check, and the version-bump step runs inside that
quality job — so there is no push path into main that skips it. The single command
that runs the whole battery locally is in AGENTS.md under "Commands".
Main is protected: a PR with a green quality check is the only way in, admins
included — the ruleset lists no bypass actors.
Everything else here is a prompt, and prompts have no typechecker — they are
reviewed against docs/prompt-standards.md, this repo's own
standard and the same checklist it scaffolds into every project. A prompt-quality defect
here is a defect in the shipped product.
Repo layout and the two non-obvious design decisions:
docs/architecture.md.