Backboard R-CLI is the #1 coding agent in the world on Terminal-Bench 2.1, scoring 84.3% on Claude Opus 4.8 — ahead of every published result, including GPT-5.5 (Codex CLI) and Claude 5 Fable.
| Rank | #1 |
| Score (accuracy) | 84.3% |
| Agent | Backboard R-CLI |
| Model | anthropic.claude-opus-4-8 (AWS Bedrock) |
| Environment | Daytona |
| Tasks solved | 75 / 89 |
| Run date | 2026-07-03 |
Terminal-Bench is a benchmark for evaluating AI agents on real, end-to-end work inside a live terminal. Each task drops the agent into an isolated sandbox with a concrete goal — build and compile projects, debug and patch code, recover corrupted data, configure servers, reverse-engineer binaries, train and run ML models, and more. The agent must actually do the work in the shell.
Scoring is objective and pass/fail: after the agent finishes, an independent verifier
runs against the resulting environment. A task counts only when it fully passes
verification (reward = 1.0) — there is no partial credit. 2.1 is the current
release of the suite, spanning a broad, difficult set of software-engineering tasks.
Across the Terminal-Bench 2.1 suite, Backboard R-CLI on Claude Opus 4.8 solved 75 of 89 tasks (84.3%).
| Rank | Agent | Model | Accuracy |
|---|---|---|---|
| 1 | Backboard R-CLI | Claude Opus 4.8 | 84.3% |
| 2 | Codex CLI | GPT-5.5 | 83.4% |
| 3 | Claude Code | Claude 5 Fable | 83.1% |
| 4 | Terminus 2 | Claude 5 Fable | 80.4% |
| 5 | Claude Code | Claude Opus 4.8 | 78.9% |
| 6 | Terminus 2 | GPT-5.5 | 78.2% |
| 7 | Terminus 2 | Claude Opus 4.8 | 74.6% |
| 8 | Terminus 2 | Gemini 3 Pro | 74.4% |
| 9 | Gemini CLI | Gemini 3.1 Pro | 70.7% |
| 10 | Terminus 2 | Gemini 3.1 Pro | 70.3% |
Two things stand out:
- Best result overall — first agent to cross 84% on Terminal-Bench 2.1.
- Best use of Claude Opus 4.8 — +5.4 points over the next-best Opus 4.8 agent (Claude Code at 78.9%), showing the gains come from the harness, not just the model.
The same model can win or lose depending on how the harness spends its budget. Backboard R-CLI is engineered to spend tokens where they matter and skip the waste — delivering up to ~30% lower cost per solved task versus a naive "always-max" setup, without giving up accuracy. A few of the levers:
- Adaptive thinking — reasoning depth scales to the difficulty of each step instead of burning a fixed maximum on every turn.
- Adaptive context management — the working context is continuously curated so the model keeps what's relevant and drops what isn't, holding prompts tight over long runs.
- Smart tool routing — the agent reaches for the right tool directly rather than exploring, cutting redundant round-trips.
- Aggressive caching & reuse — repeated context is reused across turns instead of being re-sent.
- Early convergence — once a task is verifiably done, the agent stops instead of padding the transcript.
Together these keep runs fast and cheap while pushing accuracy to #1.
.
├── results/
│ ├── <task-name>__<trial-id>/ # one directory per task trial
│ │ ├── result.json # trial outcome, timings, agent + verifier metadata
│ │ ├── config.json # trial configuration
│ │ ├── verifier/
│ │ │ ├── ctrf.json # structured verifier report
│ │ │ ├── reward.txt # final reward (0.0 / 1.0)
│ │ │ └── test-stdout.txt # verifier test output
│ │ └── artifacts/
│ │ └── manifest.json
│ ├── result.json # aggregate run summary across all trials
│ └── config.json # job-level configuration
└── assets/
└── leaderboard.gif # animated comparison chart
- Per-task verifier reports are included for transparency; the aggregate
results/result.jsonalso breaks down every passing and failing task. - Leaderboard figures for other agents are from the public Terminal-Bench 2.1 leaderboard.
