Skip to content
View theruviparambil's full-sized avatar

Block or report theruviparambil

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. veriva-eval veriva-eval Public

    Cross-model LLM-as-judge eval harness: validate AI judges with Fleiss' kappa / Krippendorff's alpha, not accuracy. Ships a real 7-model panel (Claude, GPT, Gemini, Grok, Qwen, DeepSeek, GLM) you ca…

    TypeScript 1

  2. deepeval deepeval Public

    Forked from confident-ai/deepeval

    The LLM Evaluation Framework

    Python

  3. brrr brrr Public

    Forked from anteriorcore/brrr

    high performance workflow scheduling

    Python

  4. big-finance-benchmark big-finance-benchmark Public

    Forked from Rogo-Technologies/big-finance-benchmark

    Reference harness for the Big Finance benchmark of workflow-grounded financial-research questions

    Python

  5. promptfoo promptfoo Public

    Forked from promptfoo/promptfoo

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command li…

    TypeScript

  6. ramp-analyst-evals ramp-analyst-evals Public

    An agentic finance analyst on Ramp's public agent-tool surface, plus the eval harness that grades it. Two frontier models over 22 questions x 3 samples, cross-family judging, and receipts fingerpri…

    TypeScript