slop-forensics. Jupyter Notebook
362antislop-sampler. Python
350auto-antislop. Python
164diplobench. Benchmark for LLMs playing full press diplomacy
65spiral-bench. HTML
53antislop-vllm. Jupyter Notebook
34slop-score. HTML
29lm-evaluation-harness. A framework for few-shot evaluation of language models.
5Ollama-MMLU-Pro-IRT. Ollama-MMLU-Pro fork, using a smaller IRT-tuned subset of MMLU-Pro
2entropix-gsm8k-eval. Jupyter Notebook
1not-x-but-y-bench. Python
1