This is your work, valued
LLM practitioner specializing in model training and infrastructure inference.
MINI_LLM. This is a repository used by individuals to experiment and reproduce the pre-training process of LLM.
★ 503infini-mini-transformer. This is a personal reimplementation of Google's Infini-transformer, utilizing a small 2b model. The project includes both model and training code.
★ 59MiniCharacterLLM. 这是一个一键让小参数大模型进行角色扮演的项目,从数据构成和训练都包含在这项目中
★ 27llm_learning. 这是个人关于llm的学习归纳,把我学到的知识分享给大家,让大家也能系统和简单学习llm理论及应用实战
★ 8python. Python学习
★ 1BaldEagle. Unofficial implementation of EAGLE Speculative Decoding
★ 1turbo-fieldfare. Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
★ 2.2kseta. 💻 SETA: Scaling Environments for Terminal Agents
★ 130grok-build. SpaceXAI's coding agent harness and TUI. Fullscreen, mouse interactive, extensible.
★ 24koh-my-pi. ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
★ 21klabs-molt. Python
★ 773SWE-smith. [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents
★ 721ResearchClawBench. 🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
★ 228openscience. The open-source AI workbench for scientific research
★ 3kAutomodel. 🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
★ 785MegaDLMs. GPU-optimized framework for training diffusion language models at any scale. The backend of Quokka, Super Data Learners, and OpenMoE 2 training.
★ 343orbit. Stable and Efficient Reinforcement Learning for Trillion-Parameter LLMs
★ 150speculators. A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
★ 673lucebox. LLM speculative inference server for consumer hardware & heterogeneous computing
★ 2.7kRL_Envs_101. Building and Scaling RL environments in the age of LLMs
★ 161warp. Warp is an agentic development environment, born out of the terminal.
★ 64kSWE-agent. SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]
★ 20kReagent. Agent-RRM: Exploring Reasoning Reward Model for Agents
★ 70Open-dLLM. Open diffusion language model for code generation — releasing pretraining, evaluation, inference, and checkpoints.
★ 644OpenMythos. A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
★ 15kminimax_search. MiniMax Search is an MCP (Model Context Protocol) server that provides web search and browsing capabilities.
★ 56SkillClaw. Let Skills Evolve Collectively with Agentic Evolver
★ 2.3khermes-agent. The agent that grows with you
★ 223kSWE-Gym. Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]
★ 713cuLA. CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.
★ 535Trinity-RFT. Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).
★ 673polarquant-kv. LLM KV Cache compression - K+V dual compression, 73-99% VRAM savings, zero accuracy loss
★ 57ToolRL. Python
★ 513verl-tool. A version of verl to support diverse tool use [TMLR 2026]
★ 1kA-mem. A-MEM: Agentic Memory for LLM Agents
★ 1.1kCoMLRL. Open-Source Library for Fully Cooperative Multi-LLM Reinforcement Learning
★ 110MARTI. A Framework for LLM-based Multi-Agent Reinforced Training and Inference
★ 540oh-my-codex. OmX - Oh My codeX: Your codex is not alone. Add hooks, agent teams, HUDs, and so much more.
★ 32kclaw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195kPostTrainBench. Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
★ 478pi. AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
★ 81ksglang-omni. SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
★ 716supermemory. Memory and context engine + app that is extremely fast, scalable, and can be run fully locally. The Memory API for the AI era.
★ 29klsbot. Lean & Secure Bot
★ 407claw-eval. Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
★ 738MetaClaw. 🦞 Just talk to your agent — it learns and EVOLVES 🧬.
★ 3.5kAttention-Residuals.
★ 3.4kclash-for-linux-install. 😼 优雅地使用基于 clash/mihomo 的代理环境
★ 14kOpenShell. OpenShell is the safe, private runtime for autonomous AI agents.
★ 7.9kNemoClaw. Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
★ 22kTool-Star. 🔧Tool-Star: Empowering LLM-brained Multi-Tool Reasoner via Reinforcement Learning
★ 405MiroFlow. 🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI
★ 3.1kCrossWOZ. A Large-Scale Chinese Cross-Domain Task-Oriented Dialogue Dataset
★ 722cerul. Cerul — where video becomes citable. Search video by meaning — across speech, visuals, and on-screen text.
★ 156CSL. [COLING 2022] CSL: A Large-scale Chinese Scientific Literature Dataset 中文科学文献数据集
★ 673autoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93kLightAgent. LightAgent: Lightweight Python framework for OpenAI-compatible agents with tools, memory, guardrails, tracing, lifecycle hooks, multi-agent collaboration, and workflows.
★ 1.2kslime-agentic. A project implementing various agentic RL based on the Slime post-training framework
★ 511OpenClaw-RL. OpenClaw-RL: Train any agent simply by talking
★ 5.6kAcontext. Agent Skills as a Memory Layer
★ 3.6kCLUEDatasetSearch. 搜索所有中文NLP数据集,附常用英文NLP数据集
★ 4.5kml-engineering. Machine Learning Engineering Open Book
★ 18kSelf-Evolving-Agents.
★ 1.3kskills. Give your agents the power of the Hugging Face ecosystem
★ 11knanoclaw. A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and other messaging apps,, has memory, scheduled jobs, and runs directly on Anthropic's Agents SDK
★ 30kljg-explain-concept. Claude Code skill: 8-dimensional concept anatomist. Deconstructs any concept into epiphany.
★ 212SDPO. Reinforcement Learning via Self-Distillation (SDPO)
★ 1kms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kruflo. 🌊 The leading agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
★ 67kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kQeRL. [ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
★ 512Megatron-Bridge. Training library for Megatron-based models with bidirectional Hugging Face conversion capability
★ 837SkyRL. SkyRL: A Modular Full-stack RL Library for LLMs
★ 2.1kms-agent. MS-Agent: a lightweight framework to empower agentic execution of complex tasks
★ 4.3kagentscope. Build and run agents you can see, understand and trust.
★ 28ktvm-ffi. Open ABI and FFI for Machine Learning Systems
★ 439learn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73kSkill_Seekers. Convert documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection
★ 15kFast-dLLM. Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
★ 1.1khello-agents. 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
★ 70kAionUi. Free, local, open-source 24/7 Cowork app for OpenClaw, Hermes Agent, Claude Code, Codex, OpenCode, Gemini CLI and 20+ more CLI | Customize your assistants | Star if you like it!
★ 31khpc-ops. High Performance LLM Inference Operator Library
★ 1.1kdrzero. Dr. Zero Self-Evolving Search Agents without Training Data
★ 525CLUECorpus2020. Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料
★ 1kskillsbench. SkillsBench evaluates how well skills work and how effective agents are at using them.
★ 1.6kToRL. Python
★ 353eigent. Eigent: The Open Source Cowork Desktop - Local and Free Alternative to Claude Cowork and Codex
★ 15kcc-switch. A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io
★ 123kopencode. The open source coding agent.
★ 191kEngram. Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
★ 4.6kSimpleMem. SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
★ 3.7kawesome-LLM-driven-kernel-generation. Review automated kernel generation in the era of LLMs
★ 277AgentFlow. AgentFlow: In-the-Flow Agentic System Optimization
★ 2kdflash. DFlash: Block Diffusion for Flash Speculative Decoding
★ 5.6kMiroThinker. MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
★ 8.4kwtian. Vue
★ 21VeOmni. VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
★ 2.1kR1-Searcher. R1-searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
★ 720OpenTinker. OpenTinker is an RL-as-a-Service infrastructure for foundation models
★ 677Search-R1. Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
★ 5.2kmini-sglang. A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
★ 4.7kA-mem-sys. A-MEM: Agentic Memory for LLM Agents
★ 376Scripts-SGLang. Shell
★ 7dInfer. dInfer: An Efficient Inference Framework for Diffusion Language Models
★ 476LLaDA2.X. LLaDA2.0 is the diffusion language model series developed by InclusionAI team, Ant Group.
★ 492dllm. dLLM: Simple Diffusion Language Modeling
★ 2.7kMemAgent. A MemAgent framework that can be extrapolated to 3.5M, along with a training framework for RL training of any agent workflow.
★ 1.1kAwesome-Interleaving-Reasoning. Interleaving Reasoning: Next-Generation Reasoning Systems for AGI
★ 280FlexKV. Python
★ 310cutlass. CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10kvllm-omni. A framework for efficient model inference with omni-modality models
★ 5.8kSuper-Experts-Profilling. (ICLR 2026) Unveiling Super Experts in Mixture-of-Experts Large Language Models
★ 44flashinfer. FlashInfer: Kernel Library for LLM Serving
★ 6.1kprime-rl. Agentic RL Training at Scale
★ 1.8kToolOrchestra. ToolOrchestra is an end-to-end RL training framework for orchestrating tools and agentic workflows.
★ 753auto-round. A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
★ 1.5kmiles. Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
★ 1.8kreap. REAP: Router-weighted Expert Activation Pruning for SMoE compression
★ 465Awesome-ML-SYS-Tutorial. My learning notes for ML SYS.
★ 6.8kInfinityStar. [NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
★ 774tinker-cookbook. Post-training with Tinker
★ 4knano-PEARL. Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.
★ 212prism-research. Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.
★ 72kvcached. Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
★ 1.1kagent-lightning. The absolute trainer to light up AI agents.
★ 17knanochat. The best ChatGPT that $100 can buy.
★ 57ktensorrtllm_backend. The Triton TensorRT-LLM Backend
★ 939Speech. A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
★ 18kTensorRT-LLM. TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
★ 14kModel-Optimizer. A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
★ 3.3kllm-compressor. Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
★ 3.6ktilelang. Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
★ 7.1kai-agents-for-beginners. 18 Lessons to Get Started Building AI Agents
★ 71kaibrix. Cost-efficient and pluggable Infrastructure components for GenAI inference
★ 5kLeetCUDA. LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
★ 12kflash-attention-minimal. Flash Attention in ~100 lines of CUDA (forward pass only)
★ 1.2kai-goofish-monitor. 基于 Playwright 和AI实现的闲鱼多任务实时/定时监控与智能分析系统,配备了功能完善的后台管理UI。帮助用户从闲鱼海量商品中,找到心仪产品。
★ 14kFlash-RL. Implementation for FP8/INT8 Rollout for RL training without performence drop.
★ 307ultrascale-playbook-zh. UltraScale Playbook 中文版
★ 168pansou. PanSou是一款高性能的网盘资源搜索API服务,支持TG频道和插件搜索。系统设计以性能和可扩展性为核心,支持多频道多插件并发搜索、结果智能排序和网盘类型分类。docker集成前后端,一键启动,开箱即用。仅供学习研究,请勿以各种形式用于盈利目的。
★ 14kLMCache. LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 11kfocalboard. Focalboard is an open source, self-hosted alternative to Trello, Notion, and Asana.
★ 26kMegatron-LM. Ongoing research training transformer models at scale
★ 17kslime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kSpecForge. Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
★ 1kTriton-distributed. Distributed Compiler based on Triton for Parallel Systems
★ 1.5knn-zero-to-hero. Neural Networks: Zero to Hero
★ 24kcrawl4ai. 🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
★ 76kMetis-RISE. Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
★ 22Mac-list. Mac软件清单、Mac使用技巧整理,正在不断完善中。努力做到最全。
★ 4.4kgemlite. Fast low-bit matmul kernels in Triton
★ 479how-to-optim-algorithm-in-cuda. how to optimize some algorithm in cuda.
★ 3.2ksgl-learning-materials. Materials for learning SGLang
★ 862clash-verge-rev. A modern GUI client based on Tauri, designed to run in Windows, macOS and Linux for tailored proxy experience
★ 135kLLM-RL-Visualized. 🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )
★ 4.7kbuild-your-own-x. Master programming by recreating your favorite technologies from scratch.
★ 533kebooks. 收藏的一些经典的历史、政治、心理、哲学、数学、计算机方面电子书(约10万本)
★ 7.3kART. Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
★ 11kOpenBB. Open Data Platform for analysts, quants and AI agents.
★ 71kRL2. Python
★ 1.3kKimi-K2. Kimi K2 is the large language model series developed by Moonshot AI team
★ 11kTriton-Puzzles. Puzzles for learning Triton
★ 2.5knanoVLM. The simplest, fastest repository for training/finetuning small-sized VLMs.
★ 5kFlagGems. FlagGems is an operator library for large language models implemented in the Triton Language.
★ 1.1kGenAI_Agents. 50+ tutorials and implementations for Generative AI Agent techniques, from basic conversational bots to complex multi-agent systems.
★ 24kCS-Books. 🔥🔥超过1000本的计算机经典书籍、个人笔记资料以及本人在各平台发表文章中所涉及的资源等。书籍资源包括C/C++、Java、Python、Go语言、数据结构与算法、操作系统、后端架构、计算机系统知识、数据库、计算机网络、设计模式、前端、汇编以及校招社招各种面经~
★ 27kprompt-eng-interactive-tutorial. Anthropic's Interactive Prompt Engineering Tutorial
★ 37k12-factor-agents. What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?
★ 25kDeepResearch. Tongyi Deep Research, the Leading Open-source Deep Research Agent
★ 20ktriton_tutorial. Tutorials for Triton, a language for writing gpu kernels
★ 84gemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kKsanaLLM. C++
★ 546FlagScale. FlagScale is a large model toolkit based on open-sourced projects.
★ 529ai-hedge-fund. An AI Hedge Fund Team
★ 62ktext-to-lora. Hypernetworks that adapt LLMs for specific benchmark tasks using only textual task description as the input
★ 1.3knano-vllm. Nano vLLM
★ 15kROLL. An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3.3kLUFFY. Official Repository of "Learning to Reason under Off-Policy Guidance"
★ 461Absolute-Zero-Reasoner. Official Repository of Absolute Zero Reasoner
★ 1.9kIntuitor. [ICLR 2026] Learning to Reason without External Rewards
★ 421agent-distillation. Official Code Repository for the paper "Distilling LLM Agent into Small Models with Retrieval and Code Tools"
★ 250DeepEyes. Python
★ 1.3kParScale. Parallel Scaling Law for Language Model — Beyond Parameter and Inference Time Scaling
★ 480verl-agent. verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
★ 2.2kBaldEagle. 3x Faster Inference; Unofficial implementation of EAGLE Speculative Decoding
★ 85simpleRL-reason. Simple RL training for reasoning
★ 3.9kpicotron. Minimalistic 4D-parallelism distributed training framework for education purpose
★ 2.3kSeed-Thinking-v1.5.
★ 811Semi-PD. A prefill & decode disaggregated LLM serving framework with shared GPU memory and fine-grained compute isolation.
★ 127APR. [COLM 2025] Code for Paper: Learning Adaptive Parallel Reasoning with Language Models
★ 145Tina. [ICLR 2026] Tina: Tiny Reasoning Models via LoRA
★ 338legalbenchrag. This is the repo for the LegalBench-RAG Paper: https://arxiv.org/abs/2408.10343.
★ 197TokenSwift. [ICML 2025] |TokenSwift: Lossless Acceleration of Ultra Long Sequence Generation
★ 126speculative_thinking. Python
★ 34specreason. PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]
★ 75SpeculativeDecodingPapers. 📰 Must-read papers and blogs on Speculative Decoding ⚡️
★ 1.3kSpec-Bench. Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)
★ 402Jakiro. This repository is the official implementation of "Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE" [ACL 2026 Main Accepted]
★ 37Medusa. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8kHASS. Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)
★ 56AReaL. The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
★ 5.6kEAGLE. Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
★ 2.5kParallelSpeculativeDecoding. [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Length
★ 170self-speculative-decoding. Code associated with the paper **Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding**
★ 230dynamo. A Datacenter Scale Distributed Inference Serving Framework
★ 7.6kFlashMLA. FlashMLA: Efficient Multi-head Latent Attention Kernels
★ 13k