A from-scratch PyTorch implementation of TurboQuant (ICLR 2026), Google's two-stage vector quantization algorithm for compressing LLM key-value caches — enhanced with a comprehensive, research-grade ...
Linux x86_64, aarch64 macOS x86_64, arm64 Windows x86_64, arm64 ; arm64 for the windows 11 and Python 3.11-3.14 combination only.