You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
abstract: "SWE-Review-Bench is a pilot benchmark for cold code-review bug finding on SWE-bench Lite. It evaluates LLM-based and static reviewers under a no-leakage input policy: the reviewer sees only the pre-fix file content, and the oracle (the fix patch's line ranges) is read solely by the scorer. The v0.1 release covers a 20-instance pilot with three reviewers, a prompt-variant probe, and a pytest leakage suite."