Why this matters
NEJM AI, npj Digital Medicine, and JAMA AI now require a demographic / bias audit as a publication prerequisite for any medical-LLM evaluation paper. Submitting without this guarantees a desk reject from those venues, and even non-medical venues are increasingly asking for it.
Scope
- Stratify per-tag scores for base vs LoRA across the 5-seed by 200-sample holdout.
- Flag any tag where the LoRA performance shift differs from the overall shift by more than 2 sigma.
- If sample size per tag is too small for a meaningful statistical test, state that explicitly as a limitation in the paper rather than glossing over it.
CRITICAL OWNER NOTE
Peter and Sahil: the interpretation of demographic / bias signals in HealthBench tags requires clinical context. A tag-level performance gap might be a true bias signal or might reflect rubric specifics that only a clinician can assess. Please defer the interpretation portion to Zineb, Ash Doulla, and Hillary (doctors on the project). Add their GitHub handles to assignees once known. Felipe (felipeocampoos) should also be involved.
The statistical stratification (group-by, sigma computation, plotting) can be done by any engineer on the team. The interpretation of the resulting signals must be reviewed by clinicians before anything goes into the paper.
Effort
Medium.
Team allocation reminder
This work is part of the broader threats-to-validity addressed in temporal/reviewer.md Section 7.
For tasks requiring manual verification or annotation (not clinical judgment specifically): please involve Max (maximinl, already assigned), Ashley, Martha, and Sohyeon. Their GitHub handles should be added to the relevant issues by Peter/Sahil.
For tasks requiring clinical judgment (medical correctness, harm assessment, demographic context): please defer to Zineb, Ash Doulla, Hillary, and Felipe (felipeocampoos). GitHub handles for the new clinical contributors need to be added.
Why this matters
NEJM AI, npj Digital Medicine, and JAMA AI now require a demographic / bias audit as a publication prerequisite for any medical-LLM evaluation paper. Submitting without this guarantees a desk reject from those venues, and even non-medical venues are increasingly asking for it.
Scope
CRITICAL OWNER NOTE
Peter and Sahil: the interpretation of demographic / bias signals in HealthBench tags requires clinical context. A tag-level performance gap might be a true bias signal or might reflect rubric specifics that only a clinician can assess. Please defer the interpretation portion to Zineb, Ash Doulla, and Hillary (doctors on the project). Add their GitHub handles to assignees once known. Felipe (felipeocampoos) should also be involved.
The statistical stratification (group-by, sigma computation, plotting) can be done by any engineer on the team. The interpretation of the resulting signals must be reviewed by clinicians before anything goes into the paper.
Effort
Medium.
Team allocation reminder
This work is part of the broader threats-to-validity addressed in
temporal/reviewer.mdSection 7.For tasks requiring manual verification or annotation (not clinical judgment specifically): please involve Max (maximinl, already assigned), Ashley, Martha, and Sohyeon. Their GitHub handles should be added to the relevant issues by Peter/Sahil.
For tasks requiring clinical judgment (medical correctness, harm assessment, demographic context): please defer to Zineb, Ash Doulla, Hillary, and Felipe (felipeocampoos). GitHub handles for the new clinical contributors need to be added.