Skip to content

Latest commit

Β 

History

266 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

HHRI-AI LLM EvalBoard

Comprehensive LLM Benchmark Leaderboard β€” Traditional Chinese, Multilingual & General Capabilities

A live, interactive leaderboard tracking FoxBrain models against frontier baselines across Traditional Chinese, multilingual, and general-capability benchmarks.


πŸ”— Live leaderboard

πŸ‘‰ https://muhammadsaqlainaslam.github.io/tmmlu-leaderboard/


What's inside

  • Visual Leaderboard β€” sortable rankings, discipline radar map, mean accuracy per discipline bar chart with score labels, expandable per-model subject breakdowns
  • Side-by-Side Comparison β€” filterable model comparison table across all benchmark tasks
  • Light / Dark / System theme modes β€” persists across visits
  • Google-style search β€” instant filtering with match highlighting

Benchmarks covered

TMMLU+, TCEval-v2, BigBenchHard, Ο„-bench, τ²-bench, MRCR v2, AIEC β€” spanning Traditional Chinese knowledge, reasoning, agentic tool-use, and long-context retrieval.

Models tracked

FoxBrain_v1.2_70B_0630, FoxBrain_v1.5_20251203_remove_repeat_lora, NVIDIA-Nemotron-3-Super-120B-A12B-BF16, NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, gemma-4-31B-it, Gemma-3-TAIDE-12b-Chat-ZH-TW-Prompt, gpt-oss-20b

Related projects


Curated by Muhammad Saqlain Β· HHRI-AI / Foxconn AI Research Center

Last updated: June 15, 2026

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors