| repo | openclaw/clawbench |
|---|---|
| url | https://github.com/openclaw/clawbench |
| content_timestamp | 2026-06-04 |
| time_slice | 2026-06 |
| timestamp_source | web_observed_public_github_page_2026_06_04 |
| collected_at | 2026-06-04 16:00:00 +0800 |
| source | github |
GitHub - openclaw/clawbench: ClawBench is a benchmark for agent systems that scores the full stack through execution traces, reliability metrics, and diagnostics rather than only final-task success.
Source: https://github.com/openclaw/clawbench
This raw-style public GitHub page capture was refreshed by the hourly public metadata update. Shell GitHub API access remained blocked in this workspace, so freshness is web-observed rather than API-verified.
- Repository: openclaw/clawbench
- URL: https://github.com/openclaw/clawbench
- Stars: 106
- Forks: 19
- Commits: 121
- Issues: 0
- Pull requests: 2
- License: MIT
- Primary language / stack signal: Python/Trace-Scored Benchmark/Docker
- Collection timestamp: 2026-06-04T16:00:00+08:00
- The public GitHub page showed 106 stars, 19 forks, 121 commits, 0 issues, 2 pull requests, and MIT licensing.
- The benchmark explicitly scores the harness, configuration, and model stack instead of treating the LLM alone as the system.
- README sections expose execution-trace scoring, reliability quantification, variance decomposition, and partner trace specs.
- No authenticated GitHub API freshness was used in this workspace.
No benchmark was run, no source clone was modified, and no private or authenticated metadata was used. This file preserves public page evidence for downstream classification, model-card analysis, public reports, and the site index.