Popular repositories Loading
-
veriva-eval
veriva-eval PublicCross-model LLM-as-judge eval harness: validate AI judges with Fleiss' kappa / Krippendorff's alpha, not accuracy. Ships a real 7-model panel (Claude, GPT, Gemini, Grok, Qwen, DeepSeek, GLM) you ca…
TypeScript 1
-
-
-
big-finance-benchmark
big-finance-benchmark PublicForked from Rogo-Technologies/big-finance-benchmark
Reference harness for the Big Finance benchmark of workflow-grounded financial-research questions
Python
-
promptfoo
promptfoo PublicForked from promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command li…
TypeScript
-
ramp-analyst-evals
ramp-analyst-evals PublicAn agentic finance analyst on Ramp's public agent-tool surface, plus the eval harness that grades it. Two frontier models over 22 questions x 3 samples, cross-family judging, and receipts fingerpri…
TypeScript
If the problem persists, check the GitHub status page or contact support.



