From 335a00772b71a2c58c4dcbe9bfee1c40b8128386 Mon Sep 17 00:00:00 2001 From: connerlambden Date: Thu, 4 Jun 2026 23:41:00 -0600 Subject: [PATCH] Add BGPT REFUTE benchmark (scientific critique & epistemic calibration) --- README.md | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/README.md b/README.md index 9694971..64a5103 100644 --- a/README.md +++ b/README.md @@ -1317,3 +1317,8 @@ Curto, J., et al. "Awesome Responsible Artificial Intelligence." GitHub. Last mo **MLA (Modern Language Association) 9th Edition** Curto, J., et al. "Awesome Responsible Artificial Intelligence". *GitHub*, 2026, https://github.com/AthenaCore/AwesomeResponsibleAI. Accessed 28 May 2026. + + +## Benchmarks + +- [REFUTE](https://huggingface.co/datasets/BGPT-OFFICIAL/refute) — Apache-2.0 benchmark for scientific critique & epistemic calibration on recent (2025–2026) science summaries. Separates critique skill from calibrated truthfulness (falsification, limitations, overclaims, missing-evidence refusal, confidence calibration, planted-flaw detection). [Leaderboard](https://huggingface.co/spaces/BGPT-OFFICIAL/refute-leaderboard) · [Technical report](https://huggingface.co/datasets/BGPT-OFFICIAL/refute/blob/main/TECHNICAL_REPORT.md) · [Integrators](https://huggingface.co/datasets/BGPT-OFFICIAL/refute/blob/main/INTEGRATORS.md)