Unofficial FreeToken fork: Turing (RTX 20 series) support, image input over the OpenAI API, MTP speculative decoding, 64k context on an RTX 2060 6 GB, and layer-split serving over two consumer GPUs (Qwen3.8-Flash-Next with 128k context on two RTX 3060 12 GB, no NCCL). Written by Claude Fable 5.1.
-
Updated
Sep 6, 2026 - Python