Skip to content
#

responsible-ai-evaluation-tools

Here are 5 public repositories matching this topic...

This project allows data science teams to orchestrate their model governance processes. Ensuring models built for production environments are gated properly with the adequate flags pointing to data and/or model artefacts to review.

  • Updated Aug 26, 2026
  • Python

Scenario-driven framework for evaluating demographic bias in single- and multi-agent LLM workflows. Tracks bias across sequential stages, separates model non-determinism from systematic effects, and prevents agents inferring evaluation intent. REST API and MCP server.

  • Updated Jan 12, 2026
  • Python

Improve this page

Add a description, image, and links to the responsible-ai-evaluation-tools topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the responsible-ai-evaluation-tools topic, visit your repo's landing page and select "manage topics."

Learn more