high-performance-llm-inference-engine Inference server supporting continuous batching, KV cache management, speculative decoding and INT4/INT8 quantization