superpowers-evals. Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and deterministic post-checks.

github.com/prime-radiant-inc/superpowers-evals

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.