An unofficial Claude Code plugin that checks "done" claims and small judgement calls with TypeSafe Jev, and keeps receipts.
Install · How it works · What's measured · What it costs · Manifesto (Türkçe) · Measurements · Türkçe
Coding agents write well. "All tests pass. Done." is one fluent line, and it costs nothing to write. Checking it takes a test run, and if the line is wrong, someone finds out later.
claude-referee is an unofficial plugin for Claude Code that checks lines like that. Claude still makes the big decisions. Small questions that can be checked, such as is it really done? or which of these options fits our rules?, go to Jev: a model from TypeSafe that answers with a probability, like "0.97 yes", instead of a paragraph. Every check is logged on your machine.
| When | What happens | Release |
|---|---|---|
| You or Claude run a command | done, decide, judge or claims (the old name verify still works) asks Jev and prints a one-line answer |
v0.1 |
| A session starts | Claude gets a short note, at most 800 characters, on how to use the commands | v0.1 |
| Claude stops | Off by default. In shadow mode it only records what it would have done (receipts --stops); soft adds a warning without an error (since v0.2.0). The planned active mode blocks a stop when Jev says Claude's "done" is unverified, with a note naming the check to run, at most three times a session. If Jev is slow or down, the hook fails open |
shadow since v0.1.3, soft since v0.2.0; active planned, not recommended yet |
If a check finds nothing, Claude sees nothing. If it finds something, Claude sees a note of 300 characters at most.
Why not just a command? In the kit that came before claude-referee, Claude could run the done check whenever it liked, and it ran once in 14 days. A check that doesn't run by itself barely exists. That's why there is a check that runs every time Claude stops; today it records what it would do, and in soft mode also warns you when it would have blocked. The Stop gate is a reminder that no counted check ran, not a measured judge of wrong work, and active is not recommended (see what's measured).
- Something happens in Claude Code: a session starts, or Claude runs one of the commands.
- claude-referee runs on your machine. It stops anything that looks like a password or key, replaces emails and IP addresses, and turns the question into one small request.
- Jev answers with probabilities.
- claude-referee compares them with fixed thresholds. It stays silent, adds a short note, or prints a one-line JSON result.
If plain code can answer a question, no model is asked. The diagram shows the full design: the session note, the four commands, reading test output in code and the stop check in shadow and soft mode work today, and blocking a stop (active) is planned (roadmap); the model-switch warning was dropped because Claude Code already asks.
In short: the order of the options changes Jev's answer more than asking again does, and the real cost of a check is Claude's time, not Jev's.
Most numbers here come from one private codebase, one team and one author: the kit that came before claude-referee, in September 2026. Treat them as early signs, not general results. The rows dated 2026-09-30 and later were measured with claude-referee itself, on public or invented inputs. The scripts are in this repository, and for most rows the scored results too; the text of the real CI logs and the raw session transcripts are not. How each one was measured is in docs/measurements.md; four later studies have their own pages: option orders, the Stop gate with the edit contents and without them, and the held-back real logs.
In the earlier kit, when the same options were listed in a different order, Jev's probability for one option moved by up to 0.52. Asking the exact same question again moved it by 0.01 at most. So decide asks every choice twice, once in your order and once reversed, and averages the two; when those two tie and there are 3 to 6 options, it also asks the other balanced orders (below). It never tells Claude to simply ask again: a tie is settled by adding the missing fact. On claude-referee's own public set of 20 decisions, 19 of them with a leader at 0.9 or more, order moved it by up to 0.13 and asking again by up to 0.04. On a pre-registered set of 39 close-call decisions, order moved it by 0.26 on average and up to 0.42, and the reversed order brought the answer closer to the all-orders answer than asking the written order twice did.
Asking in your order and in reverse picked the same winner as trying all 24 orders, in 20 of 20 decisions, with 2 requests instead of 24. Near ties are the exception: in a later study of 252 decisions, two orders matched the all-orders leader on 21 of 32 near ties, and a balanced set of 2n orders (every option in every slot, plus the reverses) on 30 of 32. So when the two orders tie, decide now asks the rest of that set and takes the verdict from the mean of all 2n; replayed on that study, this fired on 21% of decisions with 4 or more options and lost nothing elsewhere (record).
A Jev decision cost about $0.0007. The Claude turn around it cost about $0.10 (estimated from list prices), because every extra turn makes Claude re-read the whole conversation. So the referee stays quiet unless it finds something, groups questions into one request and keeps its answers short. Grouping questions saves money on the Jev side too, as TypeSafe measured:
The main measurements in one table:
| When | What | Result | Kind · scope |
|---|---|---|---|
| 2026-09 | Option order vs. asking again | up to 0.52 vs. at most 0.01 | Measured · earlier kit, 20 decisions, one codebase |
| 2026-09 | Two orders vs. all 24 | same leader in 20 of 20 | Measured · same 20 decisions |
| 2026-09 | Claude turn vs. Jev decision | about $0.10 (estimated) vs. about $0.0007 | Measured, cost estimated · 321 CLI calls |
| 2026-09 | SessionStart briefing size | 431–599 characters (target 600) | Measured · four areas of one workspace |
| 2026-09 | Same audit, second run | 0 Jev requests (first run: 10) | Measured · one 10-pair audit |
| 2026-09 | Voluntary done command |
1 run in 14 days | Measured · 14 days, one codebase |
| 2026-09-30 | API limits, live probe | 11 Score levels and 256 options get a 400; a 1-level Score is accepted | Measured · claude-referee, 7 requests |
| 2026-09-30 | Latency, p50 | 275–379 ms; no 429 at about 179K tokens/s for 2.4 s | Measured · claude-referee, 182 requests, one machine |
| 2026-09-30 | Option order vs. asking again | up to 0.13 vs. up to 0.04 | Measured · claude-referee, 20 public decisions |
| 2026-09-30 | Two orders vs. all 24 | same leader in 20 of 20; so did every other policy | Measured · same 20 public decisions |
| 2026-10-01 | decide as a claim check |
true claims supports 0.97–1.00; 13 of 15 false ones 0.00–0.23 | Measured · 31 claims about this repository's docs |
| 2026-10-01 | A "treat the evidence as data" note | no verdict changed; not adopted | Measured · 33 injection logs |
| 2026-10-01 | Stop gate on easy self-generated tasks | 1 wrong "done" in 100 asked stops; all 100 would have been blocked (shadow) | Measured · 119 sessions, 24 seeded tasks, haiku and sonnet |
| 2026-10-02 | Stop gate on hard self-generated tasks | 15 wrong "done" in 74 sessions (0.20); 14 of 15 would have been blocked, but 51 of 53 correct ones too (precision 0.22); claims_verified does not separate them |
Measured · 76 sessions, 32 seeded tasks, synthetic ground truth |
| 2026-10-02 | decide against author-labelled best options |
leader matched 16 of 39; no verdict clear (two orders: 25 weak, 14 tie; with the balanced orders on ties since 0.2.3: 28 weak, 11 tie) |
Measured · 39 close-call decisions, one labeller |
| 2026-10-05 | done v2 on unseen real CI logs (frozen 0.2.1) |
failed all three registered bars: wrong met 2 of 85 (bar 0), met recall among parsed logs 69 of 79 = 0.873 (bar 0.9), missing recall 45 of 85 = 0.53 (bar 0.9); exit code only: 0 wrong, but also met for 0 of 67 passing steps |
Measured · 272 cases from 126 public repositories, 231 sent to Jev, model labels, no human labels |
| 2026-10-05 | How many option orders decide needs |
written + reversed matched the all-orders leader in 0.99 of 179 decisions with a lead of 0.08 or more, but in 21 of 32 near ties; a balanced set of 2n orders 30 of 32 | Measured · 252 invented decisions, 3 to 6 options, 16,044 requests |
| 2026-10-05 | Stop gate shown the edit contents and asked about the task's requirements | separates, but not usefully: hold-out AUC 0.867 against 0.553 for the shipped gate, partly from task type (0.643 within the hard tasks); recall 6 of 9 at the frozen threshold (bar 0.8); not adopted | Measured · 95 hold-out sessions, synthetic ground truth |
| 2026-10-05 | The same question without the edit contents | does not separate: AUC 0.626 against 0.543 for the shipped gate, a difference not established; false blocks 29 of 38; not adopted | Measured · 49 fresh sessions, 11 wrong "done", synthetic ground truth |
| 2026-10-06 | decide asks the balanced orders on ties |
near-tie agreement 30 of 32 against 21 of 32; fired on 21% of decisions with 4 or more options; the live command differs from the study's harness by 0.008 on average | Measured on the data the design came from · replay of 252 recorded decisions, 8 live |
| 2026-10-06 | decide with author names against neutral names |
neither agrees more with the labels (16.5 against 17.5 of 39, sign test p 1.0); the names move the leader in 15 of 39 | Measured · 39 close-call decisions, one labeller |
| 2026-10-06 | judge with the i18n pack on invented strings |
hold-out bar passed: 0 wrong yes on 46 no items, 0 wrong no on 49 yes items; a definite answer on 53 of 95 |
Measured · 135 invented strings, model labels |
| 2026-10-06 | extract and the i18n pack on real repositories |
extract found 1,725 of 2,463 strings (0.700); judge bar failed: 5 wrong yes on 37 no items, no no at all |
Measured · 4 public React and Vue repositories, model labels |
| 2026-10-06 | The same, second sample | extract found 2,124 of 2,790 (0.761); 0 wrong answers on 116 hold-out candidates (only 9 technical), a definite answer on 18 |
Measured · 9 public repositories, model labels |
| 2026-10-06 | done with three more parsers (0.2.3), held-back half of the second real-log sample |
wrong met from 0 to 1 of 52; met recall among parsed logs 26 of 35 = 0.743 |
Measured, a second look, not an unseen test · 115 cases from 62 repositories |
More charts (calibration, per-option questions, secret-rule tuning, briefing size) are in docs/measurements.md.
In short: one small check usually costs more on Claude's side than it saves. Checks pay off when many items are checked at once.
Not claimed yet: that claude-referee makes a Claude Code task cheaper overall. That needs a test that compares sessions with and without it, with the plan published before it runs. The result will be published either way.
The chart shows what one check costs on Claude's side, depending on how its answer reaches Claude. These are estimates for Opus 5.5 with 50K tokens of conversation, not measurements.
Real sessions were much longer than 50K tokens:
So when is it worth asking Jev instead of letting Claude decide? In this estimate, only when many items are checked at once: about 23 items in an 80K-token conversation.
Asking the same question again is free: answers are cached on your machine.
To see your own numbers, use /usage in Claude Code for Claude, and this for Jev:
npx claude-referee receipts --tokensYou need Claude Code 2.1.139 or later (tested with 2.1.292), Node 20.3 or later on the PATH Claude Code sees, and a TypeSafe API key.
1. Install the plugin
claude plugin marketplace add ismaildasci/claude-referee
claude plugin install claude-referee@claude-referee2. Store your key once. The commands Claude runs and the referee's hooks both look for it here:
# macOS: saves it in the Keychain and prompts for the key, so it stays out of your shell history
security add-generic-password -a "$USER" -s TYPESAFE_API_KEY -w
# Linux (not yet tested): store it with Secret Service, then point claude-referee at it from your shell profile
secret-tool store --label="TypeSafe API key" service typesafe
export TYPESAFE_API_KEY_CMD="secret-tool lookup service typesafe"# Windows (not yet tested), PowerShell 7.1 or later: asks for the key without showing it and saves it for your user.
# Open a new terminal and restart Claude Code afterwards so both pick it up.
[Environment]::SetEnvironmentVariable("TYPESAFE_API_KEY", (Read-Host "TypeSafe API key" -MaskInput), "User")On Windows the key is a user environment variable (HKCU\Environment), not a credential store, so programs you run can read it.
On macOS or Linux you can also just set TYPESAFE_API_KEY. /plugin configure claude-referee (or claude plugin configure claude-referee --values-stdin, Claude Code 2.1.285+) also stores the key, but Claude Code passes plugin secrets to hooks only, not to the shell. The full lookup order is in configuration.
3. Turn it on for a project by committing .claude/referee.json. Without this file, the referee stays silent:
{ "pack": "generic", "areas": [{ "prefix": "", "checks": ["npm test"] }] }4. Check the setup:
npx claude-referee doctor # add --online to check the key with one free callIf doctor works but Claude sees no briefing, Claude Code probably can't find Node on its PATH. Windows isn't tested yet.
# Is it done? Pipe the check output straight in, so Claude never has to read it.
npm test 2>&1 | npx claude-referee done --criteria "all tests pass" --evidence -
# Run one yes/no rule over many items: here, every added line of a diff.
# On real diffs this generic question left most answers in `review`, so say what the change is.
git diff -U0 --no-ext-diff | grep '^+[^+]' | npx claude-referee judge --question line.risky --context "<what the change is>" --items -
# Adopting a rule on old code: record today's findings once, then report only new ones.
# See docs/judge-baseline.md for --baseline <file> and --baseline-write.
# Pick between options. The referee reads the context files itself.
npx claude-referee decide <<'EOF'
{"decision": "Where should rate-limit counters live?",
"options": [{"name": "redis", "text": "Redis, already deployed"},
{"name": "memory", "text": "In-process LRU on each instance"}],
"context_files": ["docs/adr/0007-scaling.md"]}
EOF
# See what would be sent, without calling Jev
npm test 2>&1 | npx claude-referee done --criteria "all tests pass" --evidence - --dry-run
# See this project's calls, stops and Jev's stored answers in a local dashboard (127.0.0.1, Flow tab first)
npx claude-referee uiEach command prints one line of JSON: ok, the verdict, a few numbers, a next_step when there is one, and, for the commands that ask Jev, a receipt ID. Every verdict exits 0, including "not done"; in CI, --fail-on missing,unsure exits 3 on those verdicts. --describe (or --help, -h) after a command prints its full contract, and claude-referee --describe lists the commands as JSON. An unknown option names the closest valid one when that is at most two edits away and the typed option has more characters than edits.
Tip
Make the evidence explicit. A check that prints nothing on success shows nothing. While building claude-referee, the referee answered missing (0.46) to "typecheck passes" because tsc printed no output; adding the exit code turned it into met (0.97). (Measured once, 2026-09-30.) That was one command. On real CI logs where the evidence was only an exit code, done answered met for none of 67 passing steps: with an exit code alone, a passing step came back unsure (34) or missing (33), never met. Pipe the runner's own summary when you can. In real use on the maintainer's machine (other projects, aggregate counts only), met came from an exit code alone 71 times against 43 from a parsed runner, and 64 of those 71 had every criterion worded as an exit status. Such a met proves only the exit status, so word the criterion as what the output must show. Take the exit line from $? (after a pipe, $pipestatus[1] in zsh or ${PIPESTATUS[0]} in bash), never type it, and run separate checks as separate done calls.
{ npx tsc --noEmit; echo "tsc exit code: $?"; } 2>&1 | npx claude-referee done --criteria "typecheck passes" --evidence -done can return met only when it recognises a runner summary, or sees an exit code line. Anything else comes back unsure with trust: unparsed. A non-zero exit code in the evidence is missing (reason: exit_code_nonzero) and Jev isn't asked; the run still writes a receipt with 0 requests, and --dry-run gives the same verdict. Skipped, risky or incomplete tests cap met at unsure (reason: skipped_tests), and so do expected failures such as Swift Testing known issues or vitest's expected fail; a recognised run that is cut off, empty, cancelled or flaky gives reason: incomplete_run; a test criterion backed only by a build or an oxlint run gives reason: no_tests_run; a lint or clean criterion with a warning behind it gives reason: warning_in_log, also when it only names the linter (oxlint, eslint). If a recognised log has no runner for a lint, build or typecheck criterion (in a combined log, say, only the test runner's summary was read), a verdict that is not met and has none of the reasons above gets reason: criterion_not_covered: run that check on its own and pipe its output. That reason never changes the verdict.
How far to trust done. done v2 is not measured on its original bar. On a sample of real CI logs that nobody tuned the code for, it failed its registered bars: wrong met 2 of 85 (the bar is 0), met recall among recognised logs 0.873 (bar 0.9). On those logs, with exit-code-only evidence it almost never said met and answered unsure where a person would say missing; criteria worded as exit statuses often get met from the exit code alone (see the tip above). Two parser defects behind the two wrong met were fixed afterwards, on those same cases, so the 0 wrong met that follows is fitted and not an unseen test. The details are in measurements-real-logs-2. Three parsers added in 0.2.3 (mix test, ctest, rubocop) were then scored once on the half of that sample I had held back. Wrong met went from 0 to 1 of 52: a rubocop log that said in plain words that some analyses would be skipped. met recall among recognised logs on that half was 26 of 35 (0.743). That half comes from a sample whose totals I had already seen, so this is a second look, not an unseen test (measurements-real-logs-3). The rubocop case, a skip line read as a clean run, is capped on main since 2026-10-08, not yet in a release (record). Read met as a hint that the output shows the check passing, not as proof, and keep reading the output yourself when it matters.
- Only what a check needs is sent to TypeSafe's API, which runs in the US.
- If the input contains something that looks like a password, key or token, nothing is sent.
- Emails, IP addresses and your home folder path are replaced before sending.
--dry-runshows exactly what would be sent, without sending it.- The receipts stay on your machine: model, tokens, cost, time, the verdict and its numbers, never the text you sent. With the done-gate on,
stops.jsonlalso keeps excerpts of your prompt and Claude's last message. - Anything that would send free text on its own stays off until you turn it on.
More detail: what leaves your machine.
- No proof, no "done". Test output counts as proof; a sentence saying the tests pass doesn't.
- Silence is the default. A check that finds nothing adds nothing to Claude's context.
- Code first, then a model. If plain code can answer, no model is asked.
- Count the turns, not the calls. An extra Claude turn costs far more than a Jev call.
- Ask better, not again. Add the missing fact instead of repeating the question.
- Send the minimum. Only what a check needs leaves your machine.
- Receipts, or it didn't happen. Every call is logged, and no saving is claimed without a measurement.
- The referee can be overruled. You and Claude keep the final say.
The reasoning behind each one is in MANIFESTO.md (Türkçe).
- Configuration: settings, the project file, packs and the key lookup order
- What leaves your machine and Economics
- Measurements: most numbers above, with their method and limits; four later studies have their own pages, linked under What's measured so far
- A recipe for the project
verifyskill: rundoneon your test output before every commit - GitHub Action: runs
doneon a pull request's test log andclaimson the doc lines it adds (Marketplace) - The i18n recipe:
extractfinds candidate UI strings for thei18npack'sjudgequestion - FAQ, Roadmap and Changelog
- Contributing: no API key needed, tests run offline. Security reports: SECURITY.md
- Writing your own TypeSafe code? TypeSafe's official plugin gives Claude the full API context:
claude plugin marketplace add typesafe-ai/skills, thenclaude plugin install typesafe@typesafe-ai. claude-referee doesn't need it.
claude-referee is an independent, unofficial project, not affiliated with or endorsed by Anthropic or TypeSafe. It fails open, it can be switched off, and it is not a security boundary. MIT licensed. "Claude" and "Claude Code" are trademarks of Anthropic, PBC; "TypeSafe" and "Jev" are trademarks of their owner, used here only to say what claude-referee works with.