Skip to content

docs(benchmarking): Refresh Sentry corpus benchmark results - #390

Merged
dcramer merged 10 commits into
mainfrom
docs/sentry-benchmark-corpus
Jun 4, 2026
Merged

docs(benchmarking): Refresh Sentry corpus benchmark results#390
dcramer merged 10 commits into
mainfrom
docs/sentry-benchmark-corpus

Conversation

@dcramer

@dcramer dcramer commented Jun 4, 2026

Copy link
Copy Markdown
Member

Refresh the Sentry corpus benchmark docs with the latest clean comparison rows, traced run notes, and recorded cost language. The benchmark matrix now excludes failed or partial runs so the visible comparisons stay stable.

Result Validation

Adds a docs check that validates benchmark result summaries, scoring totals, corpus ID references, trace coverage, and stable comparison eligibility before building the docs.

Benchmark Results

Adds new GPT 5.5 low-effort, Opus 4.8 traced, and Sonnet 4.6 traced result artifacts, and marks superseded or partial rows so they remain recorded without affecting the stable matrix.

Verified with pnpm docs:check.

dcramer and others added 10 commits June 4, 2026 09:21
Document the Sentry vulnerability corpus, benchmark runbook, and current Pi model results. Record sanitized run summaries and target lists while keeping raw JSONL logs out of the docs pending sensitive-data review.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Document runtime-specific benchmark workspaces and artifact names. Use the current --config-path flag and add a label for Claude SDK model IDs in benchmark results.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the scored Claude SDK Sonnet 4.6 benchmark result and update the overview copy for the new run.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the scored Pi Opus 4.8 benchmark result, label it in the benchmark table, and document that duration depends on upstream provider reliability.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the Claude SDK Opus 4.8 Sentry corpus result and label the model in the benchmark run table.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Show runtime in the run label and move benchmark metadata into a full-width row. Remove duration, reasoning, and chunk columns from the table.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Move run metadata back into the run cell, mute the runtime label, and give numeric columns compact fixed widths.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the scored benchmark runs, compact the comparison tables, and document the verification, cost, token, and timing caveats used for the matrix.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Use one outcome-first sort order across the benchmark score, cost, and timing tables. Remove the row-level JSON links so the matrix stays focused on comparison data.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Record the latest stable benchmark rows and hide partial runs from the comparison matrix.

Add a docs check that validates benchmark result metadata before the page builds.

Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
@dcramer
dcramer marked this pull request as ready for review June 4, 2026 07:24
@dcramer
dcramer enabled auto-merge (squash) June 4, 2026 07:24
@dcramer
dcramer merged commit 7d1887b into main Jun 4, 2026
20 checks passed
@dcramer
dcramer deleted the docs/sentry-benchmark-corpus branch June 4, 2026 07:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant