Pinned Loading
-
cfpb-servicer-benchmarking
cfpb-servicer-benchmarking PublicWithin-category benchmarking of 391K CFPB consumer complaints — catches a product-mix confound a naive comparison misses, plus an LLM classification eval and full-census misrouting audit. Python/Du…
Python
-
Logmortem
Logmortem PublicAI-powered RCA generator for AWS incidents (CloudWatch logs + deploy history → structured postmortem via Claude API) — with an eval harness, adversarial prompt-injection fixtures, and a CI acceptan…
Python
-
Getitdone-teardown
Getitdone-teardown PublicSan Diego 311 data → decision memo, requirements (user stories + acceptance criteria), and an LLM dedup-feasibility eval graded against the city's own duplicate labels. Real dataset, real BA delive…
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

