autonomous AI agents in the real economy, interpretable model architectures, AI metacognition, robot manufacturing
minp_paper. Code Implementation, Evaluations, Documentation, Links and Resources for Min P paper
51moneybench_pt1_planning. Python
4quest. Quantitative evalUation of modErn LLM Sampling Techniques
2smollm_entropix_torch. Jupyter Notebook
1llm-benchmarks-public. Benchmarks and tooling for evaluating large language models (shared publicly)
1vllm. [FORK] Implementation of min-z sampling on VLLM
1calibration_paper. forking it to make changes
1fishing. Python
1