Is your feature request related to a problem? Please describe. WebLLM is a powerful tool for local-first AI, but it is often constrained by browser VRAM limits, especially when dealing with long-context windows. Maintaining a large Key-Value (KV) cache for long conversations or document processing significantly limits the maximum context length and slows down inference on consumer-grade hardware.
Describe the solution you'd like. I would like to see support for TurboQuant, a new quantization framework recently introduced by Google Research. TurboQuant uses two key techniques to achieve extreme compression:
- PolarQuant: Converts Cartesian coordinates to polar coordinates (radius and angles) to eliminate memory overhead from data normalization.
- QJL (Quantized Johnson-Lindenstrauss): A "1-bit trick" that uses a mathematical transform to eliminate residual errors from the first stage.
Integrating these kernels into the WebLLM/TVM Unity pipeline would allow for 3-bit or 4-bit KV cache quantization with zero to negligible accuracy loss.
Describe alternatives you've considered. Currently, WebLLM supports standard weight quantization (4-bit/3-bit), but extreme KV cache compression is still a major bottleneck for long-context performance.
Additional context According to Google Research, TurboQuant provides:
- 6x Reduction in KV cache memory footprint.
- Up to 8x Performance Speedup in computing attention logits on modern hardware (like H100s, with significant benefits expected for WebGPU).
- Data-Oblivious: It does not require dataset-specific tuning or model retraining.
References:
Is your feature request related to a problem? Please describe. WebLLM is a powerful tool for local-first AI, but it is often constrained by browser VRAM limits, especially when dealing with long-context windows. Maintaining a large Key-Value (KV) cache for long conversations or document processing significantly limits the maximum context length and slows down inference on consumer-grade hardware.
Describe the solution you'd like. I would like to see support for TurboQuant, a new quantization framework recently introduced by Google Research. TurboQuant uses two key techniques to achieve extreme compression:
Integrating these kernels into the WebLLM/TVM Unity pipeline would allow for 3-bit or 4-bit KV cache quantization with zero to negligible accuracy loss.
Describe alternatives you've considered. Currently, WebLLM supports standard weight quantization (4-bit/3-bit), but extreme KV cache compression is still a major bottleneck for long-context performance.
Additional context According to Google Research, TurboQuant provides:
References: