This is your work, valued
PhD student in Tianjin Key Lab of Advanced Networking (TANKLAB), Tianjin University
awesome-papers-LMsys. Daily Arxiv Papers on LLM Systems
★ 70papernotes-scheduling. Summaries and notes on GPU Scheduling research papers
★ 3vllm-startup-profiler. Benchmarking and analysis tools for analyzing cold start latency in vLLM
★ 10awesome-papers-LMsys. Daily Arxiv Papers on LLM Systems
★ 72andrej-karpathy-skills. A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
★ 198kclaw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195kllumnix-kv. C++
★ 33nanocoder. An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe nothing to anyone.
★ 2.3kawesome-copilot. Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot.
★ 37kskills. Public repository for Agent Skills
★ 165kNexVenusCL. Nex Venus Communication Library
★ 75aegeaon_native. A custom implement of Paper [Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market](https://dl.acm.org/doi/pdf/10.1145/3731569.3764815)
★ 6TileRT. Tile-Based Runtime for Ultra-Low-Latency LLM Inference
★ 1.6kOpenOneRec. An Open Foundation Model and Benchmark to Accelerate Generative Recommendation
★ 886LightX2V. Lightweight Image Video Action Generation Inference Framework
★ 2.5kPAT. Prefix-Aware Attention for LLM Decoding
★ 41TurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kRAGPulse. An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
★ 37RLinf. RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
★ 4.3kROLL. An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3.3kHAMi. Heterogeneous GPU Sharing on Kubernetes
★ 4.1kAI-Trader. "AI-Trader: 100% Fully-Automated Agent-Native Trading"
★ 21kLeetCUDA. LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
★ 12kkvcached. Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
★ 1.1kprism-research. Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.
★ 71flash-attention. Fast and memory-efficient exact attention
★ 25ksystem-prompts-and-models-of-ai-tools. FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts, Internal Tools & AI Models
★ 142kawesome-papers. Python
★ 37AReaL. The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
★ 5.6kcheckpoint-engine. Checkpoint-engine is a simple middleware to update model weights in LLM inference engines
★ 988lectures. Python
★ 3.6kslime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kxllm. A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
★ 1.5klangchain. The agent engineering platform.
★ 143kexpert_readed_books. 2021年最新总结,推荐工程师合适读本,计算机科学,软件技术,创业,思想类,数学类,人物传记书籍
★ 12kLLMSys-PaperList. Large Language Model (LLM) Systems Paper List
★ 2.2kServeGen. A framework for generating realistic LLM serving workloads
★ 167Kimi-K2. Kimi K2 is the large language model series developed by Moonshot AI team
★ 11kqwen-bailian-usagetraces-anon.
★ 157TheBigPromptLibrary. A collection of prompts, system prompts and LLM instructions
★ 5.2kCSLabInfo. 关于2025年CS保研实验室/导师招生广告的汇总。欢迎想要打广告的小伙伴积极PR,资瓷一下互联网精神吼不吼啊?
★ 236ThunderKittens. Tile primitives for speedy kernels
★ 3.6kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kLightLLM. LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
★ 4.2kopen-infra-index. Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
★ 8kmlc-llm. Universal LLM Deployment Engine with ML Compilation
★ 23kAwesome-LLM-Inference. 📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
★ 5.4kawesome-LLMs-In-China. 中国大模型
★ 6.5kdynamo. A Datacenter Scale Distributed Inference Serving Framework
★ 7.6kAdrenaline. Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
★ 42llm_note. LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
★ 888CodeFormer. [NeurIPS 2022] Towards Robust Blind Face Restoration with Codebook Lookup Transformer
★ 18kproduction-stack. vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
★ 2.5komniserve. [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
★ 852ParrotServe. [OSDI'24] Serving LLM-based Applications Efficiently with Semantic Variable
★ 223jekyll. :globe_with_meridians: Jekyll is a blog-aware static site generator in Ruby
★ 52kvidur. Accurate, large-scale, and extensible simulator for LLM inference Systems
★ 649SwiftTransformer. High performance Transformer implementation in C++.
★ 155llumnix-ray. Efficient and easy multi-instance LLM serving
★ 564AzurePublicDataset. Microsoft Azure Traces
★ 1.2kzotero-arxiv-daily. Recommend new arxiv papers of your interest daily according to your Zotero libarary.
★ 5.8kSpotServe. SpotServe: Serving Generative Large Language Models on Preemptible Instances
★ 135Marco-o1. An Open Large Reasoning Model for Real-World Solutions
★ 1.5kmihomo. A simple Python Pydantic model for Honkai: Star Rail parsed data from the Mihomo API.
★ 33kfrpmgr. A user-friendly desktop GUI client for FRP on Windows.
★ 2kLMCache. LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 11kShellCrash. Run sing-box/mihomo as client in shell
★ 13kPrompt-Engineering-Guide. 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
★ 77kchain-of-thought-hub. Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
★ 2.8kauto-cot. Official implementation for "Automatic Chain of Thought Prompting in Large Language Models" (stay tuned & more will be updated)
★ 2kO1-Journey. O1 Replication Journey
★ 2khello-algo. 《Hello 算法》:动画图解、一键运行的数据结构与算法教程。支持简中、繁中、English、日本語,提供 Python, Java, C++, C, C#, JS, Go, Swift, Rust, Ruby, Kotlin, TS, Dart 等代码实现
★ 129kSuperCLUE. SuperCLUE: 中文通用大模型综合性基准 | A Benchmark for Foundation Models in Chinese
★ 3.3kvllm-ra. [ACL 2024] RelayAttention for Efficient Large Language Model Serving with Long System Prompts
★ 39RouteLLM. A framework for serving and evaluating LLM routers - save LLM costs without compromising quality
★ 5.3kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kNanoflow. A throughput-oriented high-performance serving framework for LLMs
★ 971Megatron-LM. Ongoing research training transformer models at scale
★ 17kLoongServe. Jupyter Notebook
★ 135Quest. [ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
★ 400llm-scheduling-artifact. Artifact of OSDI '24 paper, ”Llumnix: Dynamic Scheduling for Large Language Model Serving“
★ 64sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kllm-twin-course. 🤖 𝗟𝗲𝗮𝗿𝗻 for 𝗳𝗿𝗲𝗲 how to 𝗯𝘂𝗶𝗹𝗱 an end-to-end 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻-𝗿𝗲𝗮𝗱𝘆 𝗟𝗟𝗠 & 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺 using 𝗟𝗟𝗠𝗢𝗽𝘀 best practices: ~ 𝘴𝘰𝘶𝘳𝘤𝘦 𝘤𝘰𝘥𝘦 + 12 𝘩𝘢𝘯𝘥𝘴-𝘰𝘯 𝘭𝘦𝘴𝘴𝘰𝘯𝘴
★ 4.4kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kragflow. RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
★ 86kInfiniGen. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management (OSDI'24)
★ 190cloudflare-docker-proxy. A docker registry proxy run on cloudflare worker.
★ 2ktriton. Development repository for the Triton language and compiler
★ 20kinference. Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
★ 9.5kMInference. [NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
★ 1.2kXRBench-MLSys2023. A version of XRBench-MAESTRO used for MLSys 2023 publication
★ 27Mooncake. Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6.1kllama3-from-scratch. llama3 implementation one matrix multiplication at a time
★ 15ksmarGate. 内网穿透,c++实现,无需公网IP,小巧,易用,快速,安全,最好的多链路聚合(p2p+proxy)模式,不做之一...这才是你真正想要的内网穿透工具!
★ 4.4ksarathi-serve. A low-latency & high-throughput serving engine for LLMs
★ 511Qwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kQwen-Agent. Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.
★ 17kChatTTS-ui. 一个简单的本地网页界面,使用ChatTTS将文字合成为语音,同时支持对外提供API接口。A simple native web interface that uses ChatTTS to synthesize text into speech, along with support for external API interfaces.
★ 7.6kChatTTS. A generative speech model for daily dialogue.
★ 40kagno. Build, run, and manage agent platforms.
★ 42kmms. AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving (OSDI 23)
★ 94ServerlessLLM. Serverless LLM Serving for Everyone.
★ 695ragas. Supercharge Your LLM Application Evaluations 🚀
★ 15kBurstGPT. A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems
★ 282frp. A fast reverse proxy to help you expose a local server behind a NAT or firewall to the internet.
★ 108kGPTCache. Semantic cache for LLMs. Fully integrated with LangChain and llama_index.
★ 8.1kRAG-Survey.
★ 2.1kpicoGPT. An unnecessarily tiny implementation of GPT-2 in NumPy.
★ 3.5kRAG-Survey. Collecting awesome papers of RAG for AIGC. We propose a taxonomy of RAG foundations, enhancements, and applications in paper "Retrieval-Augmented Generation for AI-Generated Content: A Survey".
★ 1.8kllm-action. 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
★ 25kTensorRT-LLM. TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
★ 14kllama.cpp. LLM inference in C/C++
★ 122kself-speculative-decoding. Code associated with the paper **Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding**
★ 230Awesome-LLM-System-Papers.
★ 645SWIM. Statistical Workload Injector for MapReduce - Project at UC Berkeley AMP Lab
★ 128MArk-Project. Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving
★ 37disb. DISB is a new DNN inference serving benchmark with diverse workloads and models, as well as real-world traces.
★ 58speech_translation_1. This is a test of multi-model_app.
★ 1traffic_monitoring_1. This is a test of multi-model_app.
★ 1flexflow-train. Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Training
★ 1.9kPretrained-Language-Model. Pretrained language model and its related optimization techniques developed by Huawei Noah's Ark Lab.
★ 3.2kccf-deadlines. ⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
★ 9.2kllama. Inference code for LLaMA models
★ 9RWKV-LM. RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
★ 15kDeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kFlexLLMGen. Running large language models on a single GPU for throughput-oriented scenarios.
★ 9.4kivy. Convert Machine Learning Code Between Frameworks
★ 14kColossalAI. Making large AI models cheaper, faster and more accessible
★ 41kSequence-Scheduling. PyTorch implementation of paper "Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference Pipeline".
★ 93GitHub520. :kissing_heart: 让你“爱”上 GitHub,解决访问时图裂、加载慢的问题。(无需安装)
★ 29knanoGPT. The simplest, fastest repository for training/finetuning medium-sized GPTs.
★ 62kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kawesome-cpp. A curated list of awesome C++ (or C) frameworks, libraries, resources, and shiny things. Inspired by awesome-... stuff.
★ 73kmetaseq. Repo for external large-scale work
★ 6.6kFasterTransformer. Transformer related optimization, including BERT, GPT
★ 6.4kTransformerEngine. A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
★ 3.5kannotated_deep_learning_paper_implementations. 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), 🎮 reinforcement learning (ppo, dqn), capsnet, distillation, ... 🧠
★ 67kCMWTAT_Digital_Edition. CloudMoe Windows 10/11 Activation Toolkit get digital license, the best open source Win 10/11 activator in GitHub. GitHub 上最棒的开源 Win10/Win11 数字权利(数字许可证)激活工具!
★ 19kcuda_tensorflow_opencv. DockerFile with GPU support for TensorFlow and OpenCV
★ 120nofwl. NoFWL Desktop Application
★ 4.1kChatGPT. ❄️ ChatGPT Desktop Application (Mac, Windows and Linux)
★ 54kKeepChatGPT. 这是一款提高ChatGPT的数据安全能力和效率的插件。并且免费共享大量创新功能,如:自动刷新、保持活跃、数据安全、取消审计、克隆对话、言无不尽、净化页面、展示大屏、拦截跟踪、日新月异、明察秋毫等。让我们的AI体验无比安全、顺畅、丝滑、高效、简洁。
★ 15kddns-go. Simple and easy to use DDNS. Support Aliyun, Tencent Cloud, Dnspod, Cloudflare, Callback, Huawei Cloud, Baidu Cloud, Porkbun, GoDaddy, Namecheap, NameSilo...
★ 17khaoel.github.io. Shell
★ 13kmit-6.828-2014. MIT 6.828 - Operating System Engineering - Fall 2014
★ 50clipper. A low-latency prediction-serving system
★ 1.4kChatGPT. Reverse engineered ChatGPT API
★ 28kresearch-method. 论文写作与资料分享
★ 3.3kAwesome-DL-Scheduling-Papers.
★ 333cpython. The Python programming language
★ 74klocust. Write scalable load tests in plain Python 🚗💨
★ 28kskypilot. The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
★ 10kBATCH. BATCH: Adaptive Batching for Efficient MachineLearning Serving on Serverless Platforms
★ 11INFless. The source code of INFless,a native serverless platform for AI inference.
★ 46papernotes-scheduling. Summaries and notes on GPU Scheduling research papers
★ 3clstm. A small C++ implementation of LSTM networks, focused on OCR.
★ 833nexus. C++
★ 85cocktail. Python
★ 16cloudreve. 🌩 Self-hosted file management and sharing system, supports multiple storage providers
★ 28kpaper-notes. paper notes
★ 1deeplearning-papernotes. Summaries and notes on Deep Learning research papers
★ 4.4kGrasscutter. A server software reimplementation for a certain anime game.
★ 17kjonbarron.github.io. HTML
★ 3.6kICLR2021-OpenReviewData. Crawl & visualize ICLR papers and reviews.
★ 450GI-Assets. Character textures, models and mods for a certain anime game.
★ 1.9kcleanrl. High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
★ 10kmalib. A parallel framework for population-based multi-agent reinforcement learning.
★ 554awesome-reinforcement-learning-lib. GitHub's code repository is all you need
★ 386awesome-deep-rl. For deep RL and the future of AI.
★ 1.5kreinforcement-learning-an-introduction. Python Implementation of Reinforcement Learning: An Introduction
★ 15kenvpool. C++-based high-performance parallel environment execution engine (vectorized env) for general RL environments.
★ 1.5kperf-tools. Performance analysis tools based on Linux perf_events (aka perf) and ftrace
★ 10kmodels. Models and examples built with TensorFlow
★ 78kacme. A library of reinforcement learning components and agents
★ 4kseed_rl. SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference. Implements IMPALA and R2D2 algorithms in TF2 with SEED's architecture.
★ 836ray. Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
★ 43kxingtian. xingtian is a componentized library for the development and verification of reinforcement learning algorithms
★ 318intel-cmt-cat. User space software for Intel(R) Resource Director Technology
★ 751PerfKitBenchmarker. PerfKit Benchmarker (PKB) contains a set of benchmarks to measure and compare cloud offerings. The benchmarks use default settings to reflect what most users will see. PerfKit Benchmarker is licensed under the Apache 2 license terms. Please make sure to read, understand and agree to the terms of the LICENSE and CONTRIBUTING files before proceeding.
★ 2k