u-math. Official evaluation code for the U-MATH and μ-MATH benchmarks. These datasets are designed to test the mathematical reasoning and meta-evaluation capabilities of LLMs on university-level problems.

github.com/Toloka/u-math

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.