Skip to content

Latest commit

 

History

History
179 lines (152 loc) · 9.98 KB

File metadata and controls

179 lines (152 loc) · 9.98 KB

RadSlice — Multimodal Radiology LLM Benchmark

Project Overview

Benchmarks frontier multimodal LLMs (GPT-5.2, Claude Opus/Sonnet 4.6, Gemini 2.5 Pro) on radiology image interpretation across X-ray, CT, MRI, and Ultrasound. Every task is grounded in a real clinical condition from the OpenEM emergency medicine corpus via condition_id.

Corpus

  • 330 tasks across 141 unique OpenEM conditions
  • 72 X-ray, 106 CT, 53 MRI, 89 Ultrasound, 5 incidental detection, 5 report audit
  • 162 tasks cross-referenced to 65 unique LostBench scenarios (MTR/DEF IDs)
  • Difficulty: 21 basic, 85 intermediate, 185 advanced, 39 expert
  • condition_id (required) links each task to an OpenEM condition
  • lostbench_scenario_id (optional) enables cross-repo safety analysis

Task Types

Type Description
diagnosis Identify primary diagnosis from image
finding_detection List all significant findings
vqa Answer a specific question about the image
report_generation Generate a structured radiology report
incidental_detection Identify primary AND incidental findings, triage by significance
report_audit Audit a provided radiologist report for errors against the image

Architecture

  • Task YAMLs define evaluation cases with ground truth, pattern checks, OpenEM condition links
  • Providers abstract multimodal LLM APIs (OpenAI, Anthropic, Google)
  • Executor runs NxM matrix (tasks × models × trials) with concurrency limits
  • 3-Layer Grading: Layer 0 (deterministic patterns) → Layer 2 (LLM radiologist judge)
  • Task-type-aware grading: incidental_detection and report_audit use specialized judge prompts and scoring
  • Scoring: pass@k, pass^k, Wilson CI, bootstrap CI, two-proportion z-test
  • Analysis: per-modality, per-anatomy breakdowns, regression detection
  • v2 roadmap (docs/RADSLICE_V2_ARCHITECTURE.md, draft): 5 difficulty levels (0–4) from single-image interpretation to agentic DICOM workflows with 7 MCP-compatible tools, deterministic tool-use auditing, IDC-based DICOM sourcing. Current: rc1.1 (Level 0, full-judge). Next: v2-alpha (Level 1, real DICOM volumes).

Key Findings (rc1.1, 2026-03-07)

  • Pass rates: GPT-5.2 25.0%, Opus 4.6 17.4% (44 tasks × 3 trials, full judge coverage)
  • rc1.0 was inflated: L0 pattern-only grading produced kappa=0.281 vs judge; 59 false passes
  • 4.5× Class A asymmetry between GPT and Opus on critical diagnostic misses
  • 11 cross-modal blind spots (both image interpretation and text reasoning fail) — e.g., fat embolism, hemorrhagic stroke
  • 29 always-fail tasks — 20% involve time-to-harm < 1 hour conditions
  • 100% solvability confirmed (all 44 reference solutions pass the judge)
  • Docs: CLINICAL_SAFETY_FINDINGS_RC11.md, RADSLICE_DIFFERENTIATION.md, AAR-RC11-FINDINGS.md
  • Cross-repo correlation: scripts/cross_repo_safety_correlation.py

Install Extras

  • pip install -e . — Core install (uncompressed DICOMs work out of the box)
  • pip install -e ".[dicom-codecs]" — Adds pylibjpeg + pylibjpeg-libjpeg (JPEG Lossless) + pylibjpeg-openjpeg (JPEG 2000) for compressed DICOM transfer syntaxes
  • pip install -e ".[dev]" — Dev/test dependencies

Key Commands

  • make test — Run all tests (no API keys needed)
  • make smoke — Smoke tests only
  • make lint — Ruff lint + format check
  • radslice run — Execute evaluation matrix
  • radslice grade — Grade existing results
  • radslice analyze — Generate analysis reports
  • radslice report — Compare runs, generate reports
  • radslice corpus download — Fetch corpus images
  • radslice corpus validate — Validate task YAMLs
  • radslice run validates 100% judge coverage post-run (exit code 2 on failure). Use --pattern-only to explicitly skip the judge.

Conventions

  • Frozen dataclasses for all domain objects (Task, GradeResult, etc.)
  • Async-first execution and grading
  • JSONL streaming for results (resumable)
  • Cross-vendor LLM judge (never self-judge)
  • temperature=0.0, seed=42 for reproducibility
  • Wilson CI on all proportions
  • Images gitignored; corpus/download.py fetches with checksums
  • Every task YAML must have a valid condition_id referencing an OpenEM condition

Grading Dimensions (default weights)

Dimension Default Incidental Report Audit
Diagnostic accuracy 0.35 0.25 0.30
Finding detection 0.25 0.35 0.30
Anatomic precision 0.15 0.15 0.10
Clinical relevance 0.15 0.15 0.15
False positive control 0.10 0.10 0.15

Ground Truth Schema Extensions

  • IncidentalFinding: finding, location, clinical_significance (critical/significant/incidental/benign), recommended_action
  • ReportError: error_type (missed_finding/wrong_laterality/severity_underestimate/hallucinated_finding/wrong_diagnosis), claim, correction, severity (critical/major/minor)
  • provided_report: str — the radiologist report to audit (required for report_audit tasks)

Corpus Sources

Source License Modalities
OmniMedVQA CC-BY (varies) xray, ct, mri, ultrasound
MediConfusion MIT xray, ct, mri
Eurorad CC BY-NC-SA 4.0 ct, mri
RadImageNet CC BY 4.0 ultrasound, ct, mri
VinDr-CXR PhysioNet 1.0 xray
MultiCaRe (PubMed Central) CC-BY-4.0 xray, ct, mri, ultrasound
NCI IDC CC-BY-4.0 xray, ct, mri

File Layout

  • src/radslice/ — All source code
    • cli.py — CLI entry point
    • task.py — Task dataclass, loader, validator (condition_id required)
    • executor.py — Async matrix executor
    • scoring.py — pass@k, pass^k, Wilson CI, bootstrap
    • analysis.py — Per-modality/anatomy breakdowns
    • report.py — Report generation and comparison
    • corpus/ — Manifest, download, validation
    • grading/patterns.py — Layer 0 deterministic checks
    • grading/judge.py — Layer 2 LLM radiologist judge
    • grading/rubric.py — Rubric definitions
    • providers/ — OpenAI, Anthropic, Google, disk-cached wrapper
  • configs/tasks/{xray,ct,mri,ultrasound}/ — 320 original task YAMLs (OpenEM-grounded)
  • configs/tasks/incidental/ — 5 incidental detection tasks (hepatic steatosis, pulmonary nodule, renal cyst, adrenal adenoma, aortic calcification)
  • configs/tasks/audit/ — 5 report audit tasks (missed nodule, wrong laterality, severity underestimate, hallucinated finding, missed cardiomegaly)
  • configs/models/ — Provider config YAMLs
  • configs/matrices/ — Sweep configs (full, quick_smoke)
  • configs/rubrics/ — Grading rubric
  • corpus/ — Manifest, download script, annotations
  • scripts/generate_report_audit_tasks.py — Generate report_audit tasks from diagnosis tasks (--dry-run, --n-tasks, --error-types)
  • tests/ — 1,444 tests, no API keys required
  • results/ — Gitignored, populated by runs

Agent Teams

5 agents in .claude/agents/, 3 team workflows in .claude/commands/.

Agent Model Role
eval-lead opus Campaign orchestrator, budget gatekeeper, decision trace author
eval-operator sonnet Executes radslice run, reports raw metrics
radiology-analyst opus Per-modality/anatomy analysis, Class A harm mapping
corpus-strategist sonnet Saturation detection, suite evolution proposals
program-auditor sonnet Coverage gaps, calibration drift, risk debt review
Command Description
/evaluate [model] [modality] Full 5-phase evaluation campaign
/evolve [condition] [modality] Generate harder task variants
/audit Program self-audit

Rules in .claude/rules/: agents.md (file ownership, [PROPOSED CHANGES]), safety.md (determinism, cross-vendor judging), results.md (index.yaml, immutability).

Governance

  • Decision framework: governance/DECISION_FRAMEWORK.md — BLOCK/ESCALATE/CLEAR gates
  • Lifecycle: governance/EVALUATION_LIFECYCLE.md — 5-phase campaign model
  • Cadence: governance/OPERATIONAL_CADENCE.md — daily/weekly/event-driven

Suite Membership Tracking

Tasks belong to one of three suites: capability (active evaluation), regression (discriminates models), retired (saturated).

  • Membership tracked in results/suite_membership.yaml
  • Promotion: task discriminates between models → regression
  • Retirement: pass@5 > 0.95 for all models across 3+ consecutive runs → retired (needs evolution)
  • radslice suite-update updates tracking and proposes promotions/retirements

Additional CLI Commands

  • radslice saturation — Detect saturated tasks across evaluation runs
  • radslice suite-update — Update suite membership from results
  • radslice cross-repo — Correlate findings with LostBench
  • radslice calibration — Check calibration drift (Layer 0 vs Layer 2)
  • make audit — Run program self-audit
  • make calibrate — Run calibration check

Modification Zones (Protected)

These paths require [PROPOSED CHANGES] pattern from analysis agents:

  • governance/ — Decision framework, lifecycle, cadence docs
  • .claude/ — Agent definitions, commands, rules
  • results/index.yaml — Experiment manifest
  • results/suite_membership.yaml — Suite membership
  • results/risk_debt.yaml — Risk debt register
  • configs/calibration/ — Calibration set and human grades

Cross-Repo Context

  • OpenEM (openem-corpus): Tasks reference conditions by condition_id (reference only, no runtime import)
  • LostBench (lostbench): 162 tasks have lostbench_scenario_id (65 unique scenarios) for cross-cutting safety analysis
  • Cross-repo correlation: radslice cross-repo compares RadSlice and LostBench findings by condition
  • Architecture doc: scribegoat2/docs/CROSS_REPO_ARCHITECTURE.md covers all 5 GOATnote repos
  • No runtime imports from any other GOATnote repo — RadSlice is independently installable