localarena. Benchmark local and hosted LLMs on your own tasks with real, live inference — Ollama, llama.cpp, LM Studio, OpenAI, OpenRouter. Deterministic + judge scoring, Bradley–Terry ratings with confidence intervals, and a standalone offline HTML report. Python CLI and JS/TS SDK, zero runtime dependencies.

github.com/maziyarpanahi/localarena

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.