Rare find

shellbench. The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.

github.com/openclaw/shellbench

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.