KIV. KV cache middleware for 1M context on 12GB VRAM. Uses K vectors as a retrieval index to fetch V on-demand from system RAM. No model modification, no retraining. Drop-in HuggingFace cache replacement.

github.com/Babyhamsta/KIV

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.