Rubric-based AI daily report evaluation harness for news selection, report quality, and agent trajectory analysis.
-
Updated
Jun 7, 2026 - Python
Rubric-based AI daily report evaluation harness for news selection, report quality, and agent trajectory analysis.
🧬 Stop guessing which LLM prompt works best. ab_explorer evolves prompts via automated genetic A/B testing: it breeds, crosses, and mutates candidates, scores them against your rubric, and selects the fittest across generations. You get a winner balanced for accuracy, cost, and latency — with a beautiful CLI that reports every step with early stop
Add a description, image, and links to the rubric-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the rubric-evaluation topic, visit your repo's landing page and select "manage topics."