Human-verified evaluation for RAG, policies, search quality, model versions, and AI agents.
electron python macos evaluation policy-evaluation human-in-the-loop ai-agents llmops llm-evaluation rag-evaluation agent-evaluation retrieval-evaluation goldset
-
Updated
Jul 16, 2026 - Python