This is your work, valued
Pos2Distill-. [EMNLP25] Official code for "POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge Distillation"
★ 38FASA-ICLR2026. [ICLR 2026] FASA: FREQUENCY-AWARE SPARSE ATTENTION
★ 19TFRKN. This is the Two-hop Factual Reasoning for Knowledge Neurons (TFRKN) dataset from the paper Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons(EMNLP2024).
★ 2Z_deepseek_dstill. Jupyter Notebook
★ 1SkillClaw. Let Skills Evolve Collectively with Agentic Evolver
★ 2.3kAutoResearchClaw. Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
★ 14kskills. Public repository for Agent Skills
★ 166kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kSkillRL. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
★ 918MemAgent. A MemAgent framework that can be extrapolated to 3.5M, along with a training framework for RL training of any agent workflow.
★ 1.1kWan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kViLoMem. ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
★ 66large-qa-datasets. A collection of large question answering datasets
★ 438AutoDeco. Python
★ 73Pos2Distill-. [EMNLP25] Official code for "POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge Distillation"
★ 38Glyph. Official Repository for "Glyph: Scaling Context Windows via Visual-Text Compression"
★ 595DeepSeek-OCR. Contexts Optical Compression
★ 24kD2O. [ICLR 2025🔥] D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
★ 27loki. Algorithms for approximate attention in LLMs
★ 22hydra. Hydra is a framework for elegantly configuring complex applications
★ 11kstreaming-llm. [ICLR 2024] Efficient Streaming Language Models with Attention Sinks
★ 7.3kQuest. [ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
★ 400ASVD4LLM. Activation-aware Singular Value Decomposition for Compressing Large Language Models
★ 92VideoRoPE. [ICML 2025 Oral] An official implementation of VideoRoPE & VideoRoPE++
★ 223LLaVA-NeXT. Python
★ 4.7kneedle-in-a-haystack. Doing simple retrieval from LLM models at various context lengths to measure accuracy
★ 2.4kAwesome-LLM-KV-Cache. Awesome-LLM-KV-Cache: A curated list of 📙Awesome LLM KV Cache Papers with Codes.
★ 461MInference. [NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
★ 1.2kWUDI-Merging. The official repository of "Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors""
★ 50Block-Attention. Python
★ 48Awesome-LLM-Long-Context-Modeling. 📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
★ 2.2kDeepEP. DeepEP: an efficient expert-parallel communication library
★ 9.9kOpenRLHF. An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
★ 9.9ktrl. Train transformer language models with reinforcement learning.
★ 19kopen-r1. Fully open reproduction of DeepSeek-R1
★ 26klost-in-the-middle. Code and data for "Lost in the Middle: How Language Models Use Long Contexts"
★ 389BIG-Bench-Hard. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
★ 568Rank-Calibration. This is the repo for constructing a comprehensive and rigorous evaluation framework for LLM calibration.
★ 14FollowRAG. The demo, code and data of FollowRAG
★ 75SaySelf. Public code repo for paper "SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales"
★ 113fMRLRec. [EMNLP 2024 Findings] Official Implementation of Full-Scale Matryoshka Representation Learning for Multimodal Recommendation (fMRLRec).
★ 13Awesome-LLM-Uncertainty-Reliability-Robustness. Awesome-LLM-Robustness: a curated list of Uncertainty, Reliability and Robustness in Large Language Models
★ 831TransformerLens. A library for mechanistic interpretability of GPT-style language models
★ 3.7kcsbench. Python
★ 46llm-hallucination-survey. Reading list of hallucination in LLMs. Check out our new survey paper: "Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models"
★ 1.1kHow-To-Think-Step-by-Step. How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
★ 26Awesome-Reasoning-Foundation-Models. ✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models
★ 655MQuAKE. [EMNLP 2023] MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
★ 124RippleEdits. Evaluating the Ripple Effects of Knowledge Editing in Language Models
★ 57LRURec. [WSDM 2024] Official PyTorch Implementation of Linear Recurrent Units for Sequential Recommendation (LRURec)
★ 70reasoning-teacher. Large Language Models Are Reasoning Teachers (ACL 2023)
★ 345char-rnn. Multi-layer Recurrent Neural Networks (LSTM, GRU, RNN) for character-level language models in Torch
★ 12kDeepLearning. 深度学习入门教程, 优秀文章, Deep Learning Tutorial
★ 18kABL-HED. Handwritten Equations Decipherment with Abductive Learning
★ 102