This is your work, valued
Randomized-YaRN. Python
★ 2VRRL. Official codebase for VRRL: Visual Grounded Self-Reflection for VLMs via RL
★ 3Dr-Claude-Code. Python
★ 5Gelato. 🍨 Gelato — From Data Curation to Reinforcement Learning: Building a Strong Grounding Model for Computer-Use Agents
★ 46SkillNet. Create, Evaluate, and Connect AI Skills
★ 1.1kskillsbench. SkillsBench evaluates how well skills work and how effective agents are at using them.
★ 1.6kDySCO. DySCO: Dynamic Attention-Scaling Decoding for Long-Context LMs
★ 17GroundCUA. GroundCUA
★ 132STAT. Skill-Targeted Adaptive Training
★ 24composable_cot. Python
★ 2openrlhf-pretrain. Code for "Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining"
★ 29darling. Official Implementation of the paper "Jointly Reinforcing Diversity and Quality in Language Model Generations"
★ 61EasySteer. A Unified Framework for High-Performance and Extensible LLM Steering
★ 288open-thoughts. Fully open data curation for reasoning models
★ 2.3kVTool-R1. [ICLR 2026] "VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use"
★ 200OrionBench. Jupyter Notebook
★ 38octotools. OctoTools: An agentic framework with extensible tools for complex reasoning
★ 1.5kChartMuseum. [NeurIPS 2025] ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
★ 24CharXiv. [NeurIPS 2024] CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
★ 160sam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kReFocus_Code. Codes for ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding [ICML 2025]]
★ 50OpenThinkIMG. OpenThinkIMG is an end-to-end open-source framework that empowers Large Vision-Language Models to think with images.
★ 123Awesome_Think_With_Images. Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
★ 1.5kQRHead. QRHead: Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
★ 40knowledge-graph-of-thoughts. Official Implementation of "Affordable AI Assistants with Knowledge Graph of Thoughts"
★ 232OpenRLHF. An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
★ 9.9kLong-to-Short-via-Model-Merging. Model merging is a highly efficient approach for long-to-short reasoning.
★ 103STAR. Repo for "STAR: Spectral Truncation and Rescale for Model Merging" [NAACL 2025]
★ 10tensor2struct-public. Semantic parsers based on encoder-decoder framework
★ 92Localize-and-Stitch. Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic
★ 32axbench. Stanford NLP Python library for benchmarking the utility of LLM interpretability methods
★ 210LLM-Adapters. Code for our EMNLP 2023 Paper: "LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models"
★ 1.2kfusion_bench. FusionBench: A Comprehensive Benchmark/Toolkit of Deep Model Fusion
★ 236AdaMerging. AdaMerging: Adaptive Model Merging for Multi-Task Learning. ICLR, 2024.
★ 114HELMET. The HELMET Benchmark
★ 221LongProc. LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
★ 36To-CoT-or-not-to-CoT. Python
★ 26lo-fit. LoFiT: Localized Fine-tuning on LLM Representations
★ 45LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kactivation-steering. [ICLR 2025] General-purpose activation steering library
★ 181synthetic_data_retrieval_heads.
★ 1OLMoE. OLMoE: Open Mixture-of-Experts Language Models
★ 1klm-evaluation-harness. A framework for few-shot evaluation of language models.
★ 13kRetrieval_Head. open-source code for paper: Retrieval Head Mechanistically Explains Long-Context Factuality
★ 241finetuning. This repository contains the code used for the experiments in the paper "Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking".
★ 32investigating-alignment. Code and data for the paper: "From Distributional to Overton Pluralism: Investigating Large Language Model Alignment"
★ 2lost-in-the-middle. Code and data for "Lost in the Middle: How Language Models Use Long Contexts"
★ 387Awesome-Attention-Heads. An awesome repository & A comprehensive survey on interpretability of LLM attention heads.
★ 412Minitron. A family of compressed models obtained via pruning and knowledge distillation
★ 384SAT-LM. SatLM: SATisfiability-Aided Language Models using Declarative Prompting (NeurIPS 2023)
★ 55Logic-LLM. The project page for "LOGIC-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning"
★ 404pal. PaL: Program-Aided Language Models (ICML 2023)
★ 525DCR. Python
★ 5CodeUpdateArena. Python
★ 17pyreft. Stanford NLP Python library for Representation Finetuning (ReFT)
★ 1.6kGSM-Plus. GSM-Plus: Data, Code, and Evaluation for Enhancing Robust Mathematical Reasoning in Math Word Problems.
★ 66tuned-lens. Tools for understanding how transformer predictions are built layer-by-layer
★ 605TOXIGEN. This repo contains the code for generating the ToxiGen dataset, published at ACL 2022.
★ 350arxiv-latex-cleaner. arXiv LaTeX Cleaner: Easily clean the LaTeX code of your paper to submit to arXiv
★ 7kMiniCheck. MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents [EMNLP 2024]
★ 216MuSR. Python
★ 57honest_llama. Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
★ 581fp-dataset-artifacts. Final project starter code for NLP classes at UT Austin: provides a HuggingFace trainer to enable studying of dataset artifacts.
★ 27LIT_auto-gen-contrast-set. Python
★ 8baukit. Python
★ 257