Skip to content

Repository files navigation

high-performance-llm-inference-engine

Inference server supporting continuous batching, KV cache management, speculative decoding and INT4/INT8 quantization

About

Inference server supporting continuous batching, KV cache management, speculative decoding and INT4/INT8 quantization

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages