Skip to content

About

No description, website, or topics provided.

Resources

Stars

14 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Backboard R-CLI — #1 on Terminal-Bench 2.1

Backboard R-CLI is the #1 coding agent in the world on Terminal-Bench 2.1, scoring 84.3% on Claude Opus 4.8 — ahead of every published result, including GPT-5.5 (Codex CLI) and Claude 5 Fable.

Terminal-Bench 2.1 leaderboard — Backboard R-CLI #1

Rank #1
Score (accuracy) 84.3%
Agent Backboard R-CLI
Model anthropic.claude-opus-4-8 (AWS Bedrock)
Environment Daytona
Tasks solved 75 / 89
Run date 2026-07-03

What is Terminal-Bench 2.1?

Terminal-Bench is a benchmark for evaluating AI agents on real, end-to-end work inside a live terminal. Each task drops the agent into an isolated sandbox with a concrete goal — build and compile projects, debug and patch code, recover corrupted data, configure servers, reverse-engineer binaries, train and run ML models, and more. The agent must actually do the work in the shell.

Scoring is objective and pass/fail: after the agent finishes, an independent verifier runs against the resulting environment. A task counts only when it fully passes verification (reward = 1.0) — there is no partial credit. 2.1 is the current release of the suite, spanning a broad, difficult set of software-engineering tasks.

Results

Across the Terminal-Bench 2.1 suite, Backboard R-CLI on Claude Opus 4.8 solved 75 of 89 tasks (84.3%).

Leaderboard comparison

Rank Agent Model Accuracy
1 Backboard R-CLI Claude Opus 4.8 84.3%
2 Codex CLI GPT-5.5 83.4%
3 Claude Code Claude 5 Fable 83.1%
4 Terminus 2 Claude 5 Fable 80.4%
5 Claude Code Claude Opus 4.8 78.9%
6 Terminus 2 GPT-5.5 78.2%
7 Terminus 2 Claude Opus 4.8 74.6%
8 Terminus 2 Gemini 3 Pro 74.4%
9 Gemini CLI Gemini 3.1 Pro 70.7%
10 Terminus 2 Gemini 3.1 Pro 70.3%

Two things stand out:

  • Best result overall — first agent to cross 84% on Terminal-Bench 2.1.
  • Best use of Claude Opus 4.8 — +5.4 points over the next-best Opus 4.8 agent (Claude Code at 78.9%), showing the gains come from the harness, not just the model.

How we did it — a leaner harness

The same model can win or lose depending on how the harness spends its budget. Backboard R-CLI is engineered to spend tokens where they matter and skip the waste — delivering up to ~30% lower cost per solved task versus a naive "always-max" setup, without giving up accuracy. A few of the levers:

  • Adaptive thinking — reasoning depth scales to the difficulty of each step instead of burning a fixed maximum on every turn.
  • Adaptive context management — the working context is continuously curated so the model keeps what's relevant and drops what isn't, holding prompts tight over long runs.
  • Smart tool routing — the agent reaches for the right tool directly rather than exploring, cutting redundant round-trips.
  • Aggressive caching & reuse — repeated context is reused across turns instead of being re-sent.
  • Early convergence — once a task is verifiably done, the agent stops instead of padding the transcript.

Together these keep runs fast and cheap while pushing accuracy to #1.

Repository layout

.
├── results/
│   ├── <task-name>__<trial-id>/     # one directory per task trial
│   │   ├── result.json              # trial outcome, timings, agent + verifier metadata
│   │   ├── config.json              # trial configuration
│   │   ├── verifier/
│   │   │   ├── ctrf.json            # structured verifier report
│   │   │   ├── reward.txt           # final reward (0.0 / 1.0)
│   │   │   └── test-stdout.txt      # verifier test output
│   │   └── artifacts/
│   │       └── manifest.json
│   ├── result.json                  # aggregate run summary across all trials
│   └── config.json                  # job-level configuration
└── assets/
    └── leaderboard.gif              # animated comparison chart

Notes

  • Per-task verifier reports are included for transparency; the aggregate results/result.json also breaks down every passing and failing task.
  • Leaderboard figures for other agents are from the public Terminal-Bench 2.1 leaderboard.

About

No description, website, or topics provided.

Resources

Stars

14 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors