Skip to content

docs(design): define DeepSWE benchmark campaign - #72

Open
fazpu wants to merge 1 commit into
mainfrom
docs/deepswe-benchmark-design
Open

docs(design): define DeepSWE benchmark campaign#72
fazpu wants to merge 1 commit into
mainfrom
docs/deepswe-benchmark-design

Conversation

@fazpu

@fazpu fazpu commented Jul 13, 2026

Copy link
Copy Markdown
Member

Summary

  • record what DeepSWE can currently score and publish, and why a multi-model loopy-loop result is an agent-system result rather than a model row
  • define four benchmark arms (M0, H1, L1, L3) that separate the fixed DeepSWE scaffold, direct team-harness, loopy-loop orchestration, and heterogeneous model routing
  • specify the Pier adapter, frozen benchmark workflow, timeout semantics, committed-patch hygiene, total usage accounting, artifact set, failure policy, and phased execution plan
  • preserve the existing D1-D7 architecture boundaries; this PR changes documentation only

Important boundary

The benchmark design is accepted for this repository, but implementation remains pending. Before a full campaign, Datacurve still needs to confirm the submission category, resource limits, repetitions, and artifact/accounting contract.

Validation

  • uv run pytest — 202 passed
  • uv run ruff check .
  • uv run pyright — 0 errors, 0 warnings
  • relative documentation links resolve
  • git diff --cached --check

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant