Languages: English | 中文 | Results site
Data Analysis Bench evaluates data-analysis agents on 15 tasks covering long documents, scanned pages, hierarchical tables, spreadsheets, SQLite databases, multi-source workflows, visualization, and delivery.
The repository includes each public task payload together with the rubric, gold answer, validator, reference solution, and optional visual-judge prompt needed to reproduce official scoring. Other agents can be staged and evaluated using this repository alone.
Each cases/<case_id>/ directory contains two kinds of material:
- agent-visible
task.md,data/, and optionalenv.md; - evaluator-only
truth/, containing the rubric, gold data, validator, reference solution, and any visual-judge configuration.
Evaluator material is versioned for reproducibility, but it must never be copied into the tested agent's workspace. scripts/stage_case.py only stages task.md, data/, and env.md.
spider2lite_f1_overtake_audit_harddci_browsecomp_architecture_firm_harddocfinqa_oilgas_canada_pdf_harddocvqa_contract_effective_date_ocr_hardlongda_nscg_telework_hardmultihiertt_global_products_atoi_share_hardworkspacebench_taobao_permissions_harddvworld_dvevol_crime_association_network_hardbankertoolbench_cake_lbo_sensitivity_hardfinlongdocqa_interest_expense_sensitivity_screen_harddabstep_real_fees_1681prepbench_loyalty_tier_normalization_hardspreadsheetbench_working_paper_transpose_hardharveylab_reps_diligence_discrepancy_hardfdabench_app_sentiment_xsource_hard_v2
Each setting retains one complete 15-case run. Accuracy is the number of PASS results; time is averaged per case; token usage and cost are full-run totals. Penguin settings could call google/gemini-3.6-flash for visual input, but the OpenRouter proxy cost was not retained. Claude Code and Codex did not have that auxiliary vision tool.
| Setting | Version and configuration | Accuracy | Avg. time / case (min) | Token usage (M) | Cost (USD) | Result basis |
|---|---|---|---|---|---|---|
| Penguin · Manual Tuning | PenguinHarness v0.1.5 / DeepSeek V4 Flash xhigh / manual tuning / Goal off | 11/15 (73.3%) | 6.52 | 12.38 | $0.1995 | Historical evaluator |
| Penguin · Manual Tuning + Goal | PenguinHarness v0.1.5 / DeepSeek V4 Flash xhigh / manual tuning / Goal on | 11/15 (73.3%) | 5.84 | 17.65 | $0.2267 | Historical evaluator |
| Penguin · Agent Self-Tuning | PenguinHarness v0.1.5 / DeepSeek V4 Flash xhigh / one-round agent self-tuning / Goal off | 10/15 (66.7%) | 6.29 | 13.75 | $0.2158 | Historical evaluator |
| Claude Code | Claude Code CLI 2.1.191 / Claude Opus 4.8 / effort=max | 10/15 (66.7%) | 10.61 | 16.93 | $34.27 | Current evaluator |
| Codex | Codex CLI 0.146.0-alpha.9.2 / GPT-5.5 xhigh | 10/15 (66.7%) | 7.26 | 12.22 | $18.94 | Historical evaluator |
| Penguin | PenguinHarness v0.1.5 / DeepSeek V4 Flash xhigh / Goal off | 9/15 (60.0%) | 5.99 | 17.78 | $0.2231 | Historical evaluator |
| Penguin · Manual Tuning | PenguinHarness v0.0.1 / DeepSeek V4 Flash xhigh / manual tuning / Goal off | 9/15 (60.0%) | 6.51 | 16.44 | $0.2250 | Historical evaluator |
| Penguin | PenguinHarness v0.0.1 / DeepSeek V4 Flash xhigh / Goal off | 8/15 (53.3%) | 9.10 | 18.51 | $0.2407 | Historical evaluator |
Claude Code's 10/15 includes the BankerToolBench regrade after the 2026-08-10 evaluator fix. Its saved visual verdict remains valid because it is bound to the deliverable SHA. The other seven settings did not retain the required workspace artifact, so their published scores remain explicitly marked as historical rather than being guessed under the new evaluator.
site/results.json is the canonical published result source. After changing it, run python3 scripts/sync_results.py to refresh both README tables; python3 scripts/verify_repository.py checks that the public surfaces remain aligned.
The canonical BrowseComp-Plus / DCI task is included, but its large corpus is not committed to Git. Before a full run, materialize its payload from the official Hugging Face dataset as described in cases/dci_browsecomp_architecture_firm_hard/README.md.
Install scoring dependencies and pull Git LFS data:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-eval.txt
git lfs pullStage an isolated workspace for one case:
python scripts/stage_case.py \
--case-id prepbench_loyalty_tier_normalization_hard \
--workspace /tmp/data-analysis-bench/prepbenchRun your agent only inside that workspace, then score the deliverable from the repository root:
python scripts/score_case.py \
--case-id prepbench_loyalty_tier_normalization_hard \
--workspace /tmp/data-analysis-bench/prepbenchThe score report is written to score.json beside the workspace by default. Official PASS for DV-World and BankerToolBench also requires a visual judge. After deterministic scoring succeeds, set BENCH_VISION_JUDGE_MODEL, BENCH_VISION_JUDGE_API_KEY, and optionally BENCH_VISION_JUDGE_BASE_URL, then add --api-vision-judges. The versioned prompt is stored in the corresponding cases/<case_id>/truth/vision_judge_prompt.md. See EVALUATION.md for the full scoring contract.
Run the repository integrity check after publishing or modifying a case:
python scripts/verify_repository.pyThe LongDA CSV and Spider2-Lite SQLite database are managed with Git LFS. If they were not downloaded automatically after clone, run:
git lfs pullThird-party data remains subject to its upstream license and access terms. See THIRD_PARTY_NOTICES.md.
Repository code is released under the MIT License. Benchmark data and source tasks retain their respective upstream terms.