delta-compress-llm. Proof of concept: Exploiting temporal coherence in LLM inference-- delta encoding for KV cache compression and weight-skip prediction. Achieves F16-quality KV cache at Q4_0 compression ratios with zero perplexity loss on llama.cpp.

github.com/cenconq25/delta-compress-llm

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.