馃殌 The feature, motivation and pitch
I started hacking on vLLM recently for RamaLama and other reasons to see what the hardware support is like compared to llama.cpp. The one big noticeable gap is Vulkan support. It would solve a bunch of problems around vLLM not running great on commodity hardware:
containers/ramalama#1677
Alternatives
Use llama.cpp
Additional context
No response
Before submitting a new issue...
馃殌 The feature, motivation and pitch
I started hacking on vLLM recently for RamaLama and other reasons to see what the hardware support is like compared to llama.cpp. The one big noticeable gap is Vulkan support. It would solve a bunch of problems around vLLM not running great on commodity hardware:
containers/ramalama#1677
Alternatives
Use llama.cpp
Additional context
No response
Before submitting a new issue...