kv-cache-profiler. Stop LLM deployments from crashing in production. Profile your model's GPU memory needs before you deploy, not after. One command shows exactly how memory scales with concurrency. Save $1000s in cloud costs and weeks of debugging time.

github.com/Siddhant-K-code/kv-cache-profiler

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.