openclaw-llm-bench. A reasoning benchmark runner for comparing LLMs as OpenClaw agents use them. 52 prompts, 3 eval sets, 11 traps, LLM-as-judge, tier-based leaderboard.

github.com/arthursoares/openclaw-llm-bench

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.