This is your work, valued
Evals @ Prime Intellect
beam. Python
bibtex-mcp. Python
Athena. Python
yet-another-applied-llm-benchmark. A benchmark to evaluate language models on questions I've previously asked them to solve.