Rare find

tool-eval-bench. Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, and structured output. Supports vLLM, SGLang, and llama.cpp.

github.com/SeraphimSerapis/tool-eval-bench

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.