A benchmark evaluating if API providers preserve LLM reasoning tokens across turns. Measures thought preservation, memory gaps, and model hallucinations/fabrications about their own past thoughts.
benchmark gpt reasoning claude llm evals openrouter deepseek llm-memory llm-benchmark llm-hallucinations reasoning-tokens api-providers thought-preservation
-
Updated
Apr 29, 2026 - Python