turboquant. First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.

github.com/OnlyTerp/turboquant

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.