One import, zero config, maximum tokens/sec. Intelligently orchestrating local LLMs and hybrid speculative decoding.
-
Updated
May 26, 2026 - Python
One import, zero config, maximum tokens/sec. Intelligently orchestrating local LLMs and hybrid speculative decoding.
To associate your repository with the fast-llm topic, visit your repo's landing page and select "manage topics."