Python toolkit for Chinese Language Understanding(CLUE) Evaluation benchmark
-
Updated
May 22, 2023 - Python
Python toolkit for Chinese Language Understanding(CLUE) Evaluation benchmark
Official repository for the "VERITE: A Robust Benchmark for Multimodal Misinformation Detection Accounting for Unimodal Bias" paper.
[COLING 2025] NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models
DINOH — a Digital, AI-assisted Infrastructure for Oral History. A multilingual, human-referenced benchmark for evaluating AI support in oral history (synthetic data, per-language scoring, WebVTT/OHMS interop).
Sovereign, agent-driven LLM evaluation harness forked from deepseek-harness (dsh) — Cordis plugin architecture, EntheAI backends, and multi-model benchmarking
A from-scratch implementation of a T5 model modified with Rotary Position Embeddings (RoPE). This project includes the code for pre-training on the C4 dataset in streaming mode with Flash Attention 2.
Hybrid Network Evaluation Protocol — multi-method evaluation for hybrid quantum-classical ML models. Classifies your quantum component as Genuine, Regularizer, Ignored, or Dead Weight with bootstrap confidence intervals.
Add a description, image, and links to the evaluation-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the evaluation-benchmark topic, visit your repo's landing page and select "manage topics."