tokenvm. TokenVM is a high-performance runtime that treats LLM KV cache and activations as a virtual memory working set across GPU VRAM → pinned host RAM → NVMe storage, with intelligent paging, prefetching, and compute-copy overlap.

github.com/Siddhant-K-code/tokenvm

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.