docs(benchmarking): Refresh Sentry corpus benchmark results - #390
Merged
Conversation
Document the Sentry vulnerability corpus, benchmark runbook, and current Pi model results. Record sanitized run summaries and target lists while keeping raw JSONL logs out of the docs pending sensitive-data review. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Document runtime-specific benchmark workspaces and artifact names. Use the current --config-path flag and add a label for Claude SDK model IDs in benchmark results. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the scored Claude SDK Sonnet 4.6 benchmark result and update the overview copy for the new run. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the scored Pi Opus 4.8 benchmark result, label it in the benchmark table, and document that duration depends on upstream provider reliability. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the Claude SDK Opus 4.8 Sentry corpus result and label the model in the benchmark run table. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Show runtime in the run label and move benchmark metadata into a full-width row. Remove duration, reasoning, and chunk columns from the table. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Move run metadata back into the run cell, mute the runtime label, and give numeric columns compact fixed widths. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Add the scored benchmark runs, compact the comparison tables, and document the verification, cost, token, and timing caveats used for the matrix. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Use one outcome-first sort order across the benchmark score, cost, and timing tables. Remove the row-level JSON links so the matrix stays focused on comparison data. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
Record the latest stable benchmark rows and hide partial runs from the comparison matrix. Add a docs check that validates benchmark result metadata before the page builds. Co-Authored-By: GPT-5 Codex <noreply@anthropic.com>
dcramer
marked this pull request as ready for review
June 4, 2026 07:24
dcramer
enabled auto-merge (squash)
June 4, 2026 07:24
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refresh the Sentry corpus benchmark docs with the latest clean comparison rows, traced run notes, and recorded cost language. The benchmark matrix now excludes failed or partial runs so the visible comparisons stay stable.
Result Validation
Adds a docs check that validates benchmark result summaries, scoring totals, corpus ID references, trace coverage, and stable comparison eligibility before building the docs.
Benchmark Results
Adds new GPT 5.5 low-effort, Opus 4.8 traced, and Sonnet 4.6 traced result artifacts, and marks superseded or partial rows so they remain recorded without affecting the stable matrix.
Verified with
pnpm docs:check.