Skip to content
#

llm-benchmark

Here are 343 public repositories matching this topic...

awesome-ai-tokenomics

A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.

  • Updated Sep 25, 2026
  • Python
gcf

The AI-native wire format for structured data. 100% comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ lossless round-trips across 17 formats. Spec v3.5.1 Stable.

  • Updated Sep 27, 2026
prism

Открытый бенчмарк LLM: какая нейросеть лучше пишет код 1С:Предприятие (BSL). Объективная оценка LLM по методике SMOP с реальным исполнением в 1С — Claude, GPT, Gemini, DeepSeek, YandexGPT, GigaChat.

  • Updated Sep 25, 2026
  • Python

Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

  • Updated Sep 23, 2026
  • JavaScript

Benchmark abierto en español de modelos de IA para negocios y agentes, con juez independiente (Phi-4). Calidad, costo, velocidad, contexto largo, trabajo agéntico y fuga de credenciales, cada uno por separado. Calculadora interactiva con tus propios pesos.

  • Updated Sep 26, 2026
  • Python

Add this topic to your repo

To associate your repository with the llm-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more