Rare find

llm-inference-bench. LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.

github.com/local-inference-lab/llm-inference-bench

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.