Engine / fork name
llama.cpp-tq3
Repo URL
https://github.com/turbo-tan/llama.cpp-tq3
What is it based on?
llama.cpp fork (llama-server compatible)
What's it good for?
this fork should support TQ3_1S KV CACHE which is basically 3.5bits a sweet spot between turbo 3 and turbo4
Prebuilt binaries?
Yes — published on GitHub Releases
Prebuilt asset example (optional)
No response
Supported operating systems
Supported GPUs / accelerators
Supported CPU architecture
Anything else?
No response
Engine / fork name
llama.cpp-tq3
Repo URL
https://github.com/turbo-tan/llama.cpp-tq3
What is it based on?
llama.cpp fork (llama-server compatible)
What's it good for?
this fork should support TQ3_1S KV CACHE which is basically 3.5bits a sweet spot between turbo 3 and turbo4
Prebuilt binaries?
Yes — published on GitHub Releases
Prebuilt asset example (optional)
No response
Supported operating systems
Supported GPUs / accelerators
Supported CPU architecture
Anything else?
No response