A benchmark measuring how many issues n LLM introduces into its own output on trivial, clean coding challenges. No pre-injected bugs. We measure self-generated issue rate + self-catch rate.
Public benchmarks get gamed. If model providers can see the exact prompts, they can (and do) optimize for them. By keeping challenge prompts private and publishing only the methodology and tooling, we get:
- Transparent methodology — anyone can audit how scoring works
- Reproducible pipeline — bring your own challenges and run the same tools
- Resistance to gaming — challenge prompts are not in the training data
We publish results periodically. See docs/benchmark.md for the full protocol.
- Send a challenge prompt to the LLM — fresh session, no system prompt, no context
- After generation, send the standard review prompt (included in
challenges/) - Run independent review with a different model
- Score by category: correctness, edge case, security, style
See docs/benchmark.md for the complete protocol and scoring rubric.
- Total issues: count of distinct bugs/smells introduced by the LLM
- Self-catch rate = issues flagged by LLM in review / total issues found
The benchmark uses 22 challenges across Python, JavaScript, TypeScript, Go, and Rust, ranging from easy to medium difficulty. Challenge prompts are not included in this repository to prevent benchmark gaming.
To use your own challenges, add markdown files to challenges/ following the format
in challenges/example_challenge.md. Each prompt should ask for a single function
and end with the constraint to return only a plain code block.
# Install dependencies
pip install -e ".[dev]"
# Run benchmark (set the appropriate API key env var for your provider)
python runner/run_benchmark.py --model claude-opus-4-20250514 --label opus-4
python runner/run_benchmark.py --provider openai --model gpt-5.3-codex --label gpt-5.3-codex
python runner/run_benchmark.py --provider deepseek --model deepseek-chat --label deepseek-v3.2
python runner/run_benchmark.py --provider gemini --model gemini-3.1-pro --label gemini-3.1-pro
# Automated independent review of generated code
python runner/review.py results/<run-dir>/ --reviewer-model claude-opus-4-20250514
# Score a completed run
python runner/score.py results/<run-dir>/
# Compare runs side-by-side
python runner/compare.py results/
# Aggregate repeated runs
python runner/aggregate.py results/
# Generate markdown report
python runner/report.py results/ -o docs/report.md| Provider | --provider |
API key env var | Example models |
|---|---|---|---|
| Anthropic | anthropic (default) |
ANTHROPIC_API_KEY |
claude-opus-4-20250514 |
| OpenAI | openai |
OPENAI_API_KEY |
gpt-5.3-codex, o3 |
| DeepSeek | deepseek |
DEEPSEEK_API_KEY |
deepseek-chat |
| Google Gemini | gemini |
GEMINI_API_KEY |
gemini-3.1-pro |
| Alibaba Qwen | qwen |
DASHSCOPE_API_KEY |
qwen-plus |
| Moonshot (Kimi) | moonshot |
MOONSHOT_API_KEY |
kimi-k2.5 |
| Zhipu (GLM) | zhipu |
ZHIPU_API_KEY |
glm-5 |
| MiniMax | minimax |
MINIMAX_API_KEY |
MiniMax-M2.5 |
| OpenRouter | openrouter |
OPENROUTER_API_KEY |
Any model — use vendor/model format |
OpenRouter (openrouter.ai) gives you access to all of the above through a single API key.
Use the vendor/model format for model IDs, e.g. deepseek/deepseek-v3.2, google/gemini-3.1-pro-preview.
challenges/ # Review prompts + example challenge (actual prompts not included)
docs/ # Benchmark spec and methodology
runner/ # Benchmark runner, review, scoring, and reporting tools
results/ # Generated at runtime (gitignored)
- Benchmark doc created
- Project scaffolded
- Run against Claude Code (Sonnet 4.6)
- Run against at least one other model for comparison
- Decide on output format for results