debate. Adversarial multi-turn benchmark for LLM debate quality, using side-swapped matchups and multi-model judging to rank models by judged debate performance.

github.com/lechmazur/debate

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.