Skip to content

Is Claude Opus 4.5 required as the judge LLM, or can cheaper alternatives be used? #1

Description

@Dereck0602

Hi, thanks for the great work on this project.

I noticed that the evaluation setup uses claude-opus-4.5 as the judge LLM. I’d like to better understand whether this model is strictly required for reliable evaluation results.

Specifically, I have a few questions:

  1. Is claude-opus-4.5 required for the judge LLM, or is it only the recommended/default option?
  2. Have you tested cheaper judge models, such as deepseek-v4-pro, for the same evaluation tasks?
  3. If a cheaper model is used as the judge LLM, is there any known degradation in evaluation quality, consistency, or correlation with human judgments?

The main motivation is cost reduction. If a cheaper model can provide comparable judging quality, it would make large-scale evaluations much more affordable.

Thanks!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions