Open-source, local-first evaluation infrastructure for applied AI systems, built for developer and agent workflows.
-
Updated
Jul 27, 2026 - Rust
Open-source, local-first evaluation infrastructure for applied AI systems, built for developer and agent workflows.
Medical AI portfolio exploring patient safety, clinical decision support, healthcare quality improvement, AI evaluation, and healthcare analytics.
Reusable audit scaffold for detecting prefill awareness confounds in transcript-based AI evals
AI Foundations evaluation of Anthropic Claude Constitution artifacts; distinguishes external behavioral source from Source of self.
Defines model weight-pressure and tests whether source-bound contact architecture can carry structure against default model collapse patterns.
Public control map for AI Foundations / Origin | Continuum evaluations, defining test categories, goals, claim boundaries, pass/fail behavior, and evidence limits.
Add a description, image, and links to the ai-evaluations topic page so that developers can more easily learn about it.
To associate your repository with the ai-evaluations topic, visit your repo's landing page and select "manage topics."