Evaluate models with a benchmark solution, "PingBench" - LLM as a judge (probably GPT-OSS-120B) - **Hammer out the rest of the details**
Evaluate models with a benchmark solution, "PingBench"