thought-kv-hierarchy. Semantics-aware memory hierarchy for KV-cache in LLM reasoning. Offloads low-importance reasoning tokens from HBM to DDR instead of evicting them, preserving accuracy at reduced memory cost.

github.com/Justin0504/thought-kv-hierarchy

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.