Benchmark harness comparing 8 Gemini inference strategies for structured generation: output quality gate, CI-gated deterministic replay benchmark, and a live results dashboard
-
Updated
Jul 17, 2026 - Python
Benchmark harness comparing 8 Gemini inference strategies for structured generation: output quality gate, CI-gated deterministic replay benchmark, and a live results dashboard
A benchmark evaluating if API providers preserve LLM reasoning tokens across turns. Measures thought preservation, memory gaps, and model hallucinations/fabrications about their own past thoughts.
independent audit layer for hidden llm reasoning token billing
Add a description, image, and links to the reasoning-tokens topic page so that developers can more easily learn about it.
To associate your repository with the reasoning-tokens topic, visit your repo's landing page and select "manage topics."