Rare find

yet-another-applied-llm-benchmark. A benchmark to evaluate language models on questions I've previously asked them to solve.

github.com/carlini/yet-another-applied-llm-benchmark

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.