Comprehensive LLM Benchmark Leaderboard β Traditional Chinese, Multilingual & General Capabilities
A live, interactive leaderboard tracking FoxBrain models against frontier baselines across Traditional Chinese, multilingual, and general-capability benchmarks.
π https://muhammadsaqlainaslam.github.io/tmmlu-leaderboard/
- Visual Leaderboard β sortable rankings, discipline radar map, mean accuracy per discipline bar chart with score labels, expandable per-model subject breakdowns
- Side-by-Side Comparison β filterable model comparison table across all benchmark tasks
- Light / Dark / System theme modes β persists across visits
- Google-style search β instant filtering with match highlighting
TMMLU+, TCEval-v2, BigBenchHard, Ο-bench, ΟΒ²-bench, MRCR v2, AIEC β spanning Traditional Chinese knowledge, reasoning, agentic tool-use, and long-context retrieval.
FoxBrain_v1.2_70B_0630, FoxBrain_v1.5_20251203_remove_repeat_lora, NVIDIA-Nemotron-3-Super-120B-A12B-BF16, NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, gemma-4-31B-it, Gemma-3-TAIDE-12b-Chat-ZH-TW-Prompt, gpt-oss-20b
- π FoxBrain EvalHub (frontier benchmark & model reference) β muhammadsaqlainaslam.github.io/foxbrain-eval-dashboard
- π FoxBrain EvalHub GitHub repo β github.com/MuhammadSaqlainAslam/foxbrain-eval-dashboard
Curated by Muhammad Saqlain Β· HHRI-AI / Foxconn AI Research Center
Last updated: June 15, 2026