This is your work, valued
NVSHMEM-Tutorial. NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer
★ 195xv6-riscv-solution. MIT 6.S081 xv6-riscv solution
★ 100hypocaust-2. hypocaust-2, a type-1 hypervisor with H extension run on RISC-V machine
★ 60hypercraft. hypercraft is a VMM library written in Rust.
★ 53hypocaust. hypocaust, a S-mode trap and emulate type-1 hypervisor run on RISC-V machine.
★ 50AttnLink. :construction: An experimental communicating attention kernel based on DeepEP.
★ 34Paper-reading. My Paper Reading Lists and Notes.
★ 25cu-x. 🎉My Collections of CUDA Kernels~
★ 11TileGraph. TileGraph is an experimental DNN compiler that utilizes static code generation and kernel fusion techniques.
★ 11TJU-CourseSharing. 天津大学课程共享计划
★ 9ixgbe-driver. Intel 82599+ 10Gb NIC Driver.
★ 6FAT32. FAT32 File System in Rust
★ 6arceos. An experimental modular OS written in Rust.
★ 5Notes. All notes of KuangjuX.
★ 5minikernel. minikernel, A minimal kernel which run on hypocaust
★ 4minigo. minigo is a minimal go compiler for compiler course
★ 4rCore-fat. 🦀️ rCore-Tutorial with fat32 file system
★ 4LeetcodeSolutions. My solution for leetcode and other algorithm problems.
★ 4TorchAlgo. Collection of algorithms implemented using PyTorch and Triton.
★ 4SimpleDB. C++
★ 3Bachelor-Thesis. Undergraduate thesis
★ 3Projects. :scroll: Collection of undergraduate homework and labs and other
★ 3flux. A fast communication-overlapping library for tensor/expert parallelism on GPUs.
★ 3who-unfollow-you. :hammer_and_wrench: A Tool to find who unfollowed you recently in github.
★ 2compiler-and-arch. A list of tutorials, paper, talks, and open-source projects for emerging compiler and architecture
★ 1skills. Skills for Real Engineers. Straight from my .agents directory.
★ 198kAwesome-LLM-Long-Context-Modeling. 📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
★ 2.2kvitamin-cuda. 🍎 One kernel a day keeps high latency away. A hands-on CUDA learning path featuring a rich collection of kernels, from the basics to peak performance, seamlessly integrated as PyTorch C++ extensions.
★ 197flash-msa. Python
★ 25MoonEP. MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
★ 960Kimi-K3. Open Frontier Intelligence
★ 7.7ktransformer-tricks. A collection of tricks and tools to speed up transformer models
★ 220AgentENV. AgentENV (AENV) is a distributed platform for running agent environments at scale.
★ 2.7kCuTe. Reference implementation and examples of the CuTe Layout representation and algebra.
★ 217DeepSpec. DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
★ 6.8kMSA. Python
★ 392auto-gpu-kernel. Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average speedup of 34.93x
★ 148inkos. Story Creation AI Agent for novel, scripts, translation, interactive games, and IP content
★ 8.6kchinese-novelist-skill. 🎭 AI 驱动的中文小说创作助手。三层递进式智能问答、跨会话偏好记忆、中断续写、每章必留悬念钩子、自动校验修复。长篇小说一次性全稿无痛完成。 ⚡ 稳定写作,推荐 非线智能 API nonelinear.com 联系客服报 [写小说] 享20元额度
★ 2.5kvibesys. Can AI Agents Build Bespoke Systems?
★ 88KernelWiki. Python
★ 319VeloQ. Agent-friendly GPU profile-query CLI
★ 107ForgeTrain. Python
★ 274kernel-design-agents.
★ 788ds4. DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
★ 20kai-video-editing-skill. AI Agent Skill for automated vlog editing. Feed raw footage, get a finished video. Powered by ffmpeg + Whisper + Vision API.
★ 91pro-video-composer. AI 视频合成 pipeline — 文稿+录音+空镜 → 自动出片。串联 ffmpeg + Remotion + ASR + agent 4 个角色。By nyx研究所
★ 32cuda-oxide. cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.
★ 3kOverleaf-Desktop. A native macOS app for syncing projects to your local filesystem
★ 68tokenspeed. TokenSpeed is a speed-of-light LLM inference engine.
★ 1.8kssa. Official repository for "SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space"
★ 27TradingAgents. TradingAgents: Multi-Agents LLM Financial Trading Framework
★ 95kKernelGYM. [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML 2026]
★ 198KSA. Kwai Summary Attention
★ 59FlashQLA. high-performance linear attention kernel library built on TileLang
★ 622pith-train. Compact and Agent-Native MoE Training System
★ 311llama.cpp-deepseek-v4-flash. Experimental implementation of DeepSeek v4 flaash in llama.cpp
★ 331LLM-RL-Visualized. 🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )
★ 4.7kpyptx. A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch
★ 370AI-Infra-Auto-Driven-SKILLS. Python
★ 708DeepPaperNote. DeepPaperNote is an agent skill for deep-reading a single paper and generating high-quality Obsidian-style research notes. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
★ 569ppt-master. AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and support for your own .pptx templates. · by Hugo He
★ 42kTileKernels. A kernel library written in tilelang
★ 1.7kzhihu-to-markdown. JavaScript
★ 9huashu-design. Huashu Design · HTML-native design skill for Claude Code · Claude Code 里 HTML 原生的设计 skill · 高保真原型 / 幻灯片 / 动画 + 20 设计哲学 + 5 维评审 + MP4 导出 · Agent-agnostic
★ 22kFlashKDA. FlashKDA: high-performance Kimi Delta Attention kernels
★ 1.1ksass-king. Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.
★ 315FM-Agent. Python
★ 443SkVM. The Language Virtual Machine for Agent Skills
★ 541html-ppt-skill. HTML PPT Studio — AgentSkill with 24 themes, 31 layouts, 20+ animations for building professional HTML presentations
★ 7.5khermes-agent. The agent that grows with you
★ 223kcuda-optimized-skill. A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It helps improve custom GPU operators with reproducible workflows and evidence-based performance comparison.
★ 194Happy-Horse-1.0. Information collection for the Happy Horse AI video generator model. Official demo and updates at happyhorses.io.
★ 636deepxiv_sdk. Talk to research papers like talking to authors - Python package with AI agent for arXiv papers
★ 762Embodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15kcuLA. CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.
★ 535gemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kptx-isa-markdown. PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.
★ 221AKO4ALL. Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language
★ 337codex. Lightweight coding agent that runs in your terminal
★ 1happyclaw. happy happy happyclaw~
★ 780Booth. Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.
★ 1.7kLongCat-Next. Python
★ 464AI-workflow.
★ 71weclaw. Connect to any agents with WeChat ClawBot.
★ 1.6kMagiCompiler. A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.
★ 324parameter-golf. Train the smallest LM you can that fits in 16MB. Best model wins!
★ 5.2kcuda-evolve-oss. Autonomous GPU kernel optimization system driven by AI agents.
★ 31SOL-ExecBench. A benchmark of real-world DL kernel problems
★ 268cutile-rs. cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
★ 722Attention-Residuals.
★ 3.4kgstack. Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
★ 126kncu-cli. Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.
★ 34PaperBanana. PaperBanana: Automating Academic Illustration For AI Scientists
★ 6.9kautokernel. Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
★ 1.5kWeFlow. WeFlow - 一个本地的微信聊天记录导出和年度报告应用
★ 13kbub. Bub it. Build it. A hook-first runtime for agents that live alongside people.
★ 1.6kSSR-V2ray-Trojan. 2026机场推荐与机场评测
★ 17kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93knsys-ai. Terminal UI for NVIDIA Nsight Systems profiles — timeline viewer, kernel navigator, NVTX hierarchy
★ 67Where2Live. Rust
★ 1Daoyou. 《万界道友》是一款以 AIGC 驱动、高自由度文字体验、修仙世界观为核心的开源游戏。在这里,你将以普通修士之身,借功法、灵根、神通、法宝与奇遇,一步步推演自己的修行之路。
★ 84infra-skills. A collection of specialized agent skills for AI infrastructure development, enabling Claude Code to write, optimize, and debug high-performance systems.
★ 140CUDA-Agent. CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
★ 1.1kcurgit. A high-performance CLI tool written in Rust that acts as a standalone Git Agent.
★ 11Humanizer-zh. Humanizer 的汉化版本,Claude Code Skills,旨在消除文本中 AI 生成的痕迹。
★ 14kawesome-claude-skills. A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows
★ 71kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385khumanizer. Agent skill that removes signs of AI-generated writing from text
★ 32kopenfang. Open-source Agent Operating System
★ 18kRepoLaunch. Automate the build, execution and test of GitHub repositories across programming languages and operating systems.
★ 129learn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73kKsanaDiT. KsanaDiT: High-Performance DiT (Diffusion Transformer) Inference Framework for Video & Image Generation
★ 62ThunderKittens-Tutorials. Makefile
★ 8claude-skills-guide. 📚 Claude Skills 开发完全指南 - 从基础到精通 | Complete guide for developing Claude Skills - from basics to mastery
★ 214pie. Pie: Programmable LLM Serving
★ 193FlyDSL. FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.
★ 257humanize. From Automated Idea Factory to Realization
★ 1.4kAI-Research-SKILLs. Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
★ 11kopeninfer. Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 619GatedDeltaNet. [ICLR 2025] Official PyTorch Implementation of Gated Delta Networks: Improving Mamba2 with Delta Rule
★ 636andrej-karpathy-skills. A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
★ 198kECC. The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
★ 237kmixture-of-experts. PyTorch Re-Implementation of "The Sparsely-Gated Mixture-of-Experts Layer" by Noam Shazeer et al. https://arxiv.org/abs/1701.06538
★ 1.2kThunderAgent. A simple, fast and robust program-aware agentic inference system.
★ 402qwen-asr. C inference for Qwen3-ASR 0.6b and 1.7b transcriptions models
★ 588flash-linear-attention. FLA but cuTile
★ 27GLM-OCR. GLM-OCR: Accurate × Fast × Comprehensive
★ 7.2kLongcat-Flash-Thinking-ZigZag. Python
★ 10AutoFigure-Edit. Python
★ 4.1kbased. Code for exploring Based models from "Simple linear attention language models balance the recall-throughput tradeoff"
★ 256Stanford-CS336. My Solution and Notes for the Stanford CS336: LLM from scratch
★ 264flashtile. FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.
★ 61sm-profiler. Python
★ 84SGLang-FluentLLM. Python
★ 114nanoclaw. A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and other messaging apps,, has memory, scheduled jobs, and runs directly on Anthropic's Agents SDK
★ 30kPaper2Slides. "Paper2Slides: From Paper to Presentation in One Click"
★ 3.8ktilelang-puzzles. Learning TileLang with 10 puzzles!
★ 355minions. Big & Small LLMs working together
★ 1.3kDeepSeek-OCR-2. Visual Causal Flow
★ 3.2kawesome-LLM-driven-kernel-generation. Review automated kernel generation in the era of LLMs
★ 278flashinfer-bench-starter-kit. FlashInfer Bench @ MLSys 2026: Building AI agents to write high performance GPU kernels
★ 178vibetensor. Our first fully AI generated deep learning system
★ 635funny_cute. Some funny cute/cuteDSL code snippets
★ 33hpc-ops. High Performance LLM Inference Operator Library
★ 1.1kbanana-slides. 一个基于nano banana pro🍌的原生AI PPT生成应用,迈向"Vibe PPT"; 支持上传任意模板图片,上传任意素材&智能解析,一句话/大纲/页面描述自动生成PPT,口头修改指定区域、一键导出可编辑ppt - An AI-native slides generator based on nano banana pro🍌
★ 15kx-algorithm. Algorithm powering the For You feed on X
★ 27kLLMSys-PaperList. Large Language Model (LLM) Systems Paper List
★ 2.2kiris.c. Flux 2 image generation model pure C inference
★ 2kMoonlight. Muon is Scalable for LLM Training
★ 1.5kGLM-Image. GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
★ 1kmhc-lite. mHC-lite: You Don’t Need 20 Sinkhorn-Knopp Iterations
★ 91flexflow-serve. FlexFlow Serve: Low-Latency, High-Performance LLM Serving
★ 87Engram. Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
★ 4.6khp_rms_norm. High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)
★ 30TensorRT-Edge-LLM. High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
★ 489sparse_attention. Examples of using sparse attention, as in "Generating Long Sequences with Sparse Transformers"
★ 1.6kcute-viz. Cute layout visualization
★ 44WeDLM. WeDLM: The fastest diffusion language model with standard causal attention and native KV cache compatibility, delivering real speedups over vLLM-optimized baselines.
★ 649sing-box. The universal proxy platform
★ 37kmHC.cu. mHC kernels implemented in CUDA
★ 265Sparse-VideoGen. [ICML2025, NeurIPS2025 Spotlight] Sparse VideoGen 1 & 2: Accelerating Video Diffusion Transformers with Sparse Attention
★ 697Triton-to-tile-IR. incubator repo for CUDA-TileIR backend
★ 151CUDA-L2. CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
★ 468cuda-tile. CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-based computation patterns and optimizations targeting NVIDIA tensor core units.
★ 1ksonic-moe. Accelerating MoE with IO and Tile-aware Optimizations
★ 734mini-sglang. A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
★ 4.7kTurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kcutile-learn. NVIDIA cuTile learn
★ 169TileGym. Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming
★ 785cutile-python. cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
★ 2.1komniserve. [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
★ 852DeepSeek-Math-V2. Python
★ 1.6kSimuMax. a static analytical model for LLM distributed training
★ 164PRISM. An Elegant Academic Homepage Builder
★ 674fastrl. [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
★ 176flex-block-attn. flex-block-attn: an efficient block sparse attention computation library
★ 129TileRT. Tile-Based Runtime for Ultra-Low-Latency LLM Inference
★ 1.6ktilelang-dsa. DeepSeek-V3.2-Exp DSA Warmup Lightning Indexer training operator based on tilelang
★ 47flash-moba. C++
★ 252EAGLE. Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
★ 2.5kFlash-Sparse-Attention. 🚀🚀 Efficient implementations of Native Sparse Attention
★ 622Kimi-Linear.
★ 1.5kminimind. 🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
★ 54kTwilight. [NeurIPS'25 Spotlight] Adaptive Attention Sparsity with Hierarchical Top-p Pruning
★ 105DeepSeek-OCR. Contexts Optical Compression
★ 24kdInfer. dInfer: An Efficient Inference Framework for Diffusion Language Models
★ 476nanochat. The best ChatGPT that $100 can buy.
★ 57kreasoning-from-scratch. Implement a reasoning LLM in PyTorch from scratch, step by step
★ 4.9kFlashMoE. Distributed MoE in a Single Kernel [NeurIPS '25]
★ 281sgl-learning-materials. Materials for learning SGLang
★ 862tiktokenizer. Online playground for OpenAPI tokenizers
★ 1.7kDeepSeek-V3.2-Exp. Python
★ 1.6kcutlass-notes. From Minimal GEMM to Everything
★ 230flash-kmeans. Fast and memory-efficient exact kmeans
★ 704Flash-RL. Implementation for FP8/INT8 Rollout for RL training without performence drop.
★ 307tiny-qwen. A minimal PyTorch re-implementation of Qwen 3.5
★ 432tokenweave. Accepted to MLSys 2026
★ 91triton-runner. Multi-Level Triton Runner supporting Python, IR, PTX, AMDGCN, cubin and hasco.
★ 99KsanaLLM. C++
★ 546Awesome-RL-for-LRMs. A Survey of Reinforcement Learning for Large Reasoning Models
★ 2.5klayout-categories. This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".
★ 140Syno. Source code repository for ASPLOS '25 paper "Syno: Structured Synthesis for Neural Operators"
★ 14chatlog. chat log tool, easily use your own chat data. 聊天记录工具,轻松使用自己的聊天数据
★ 9.2kllumnix-ray. Efficient and easy multi-instance LLM serving
★ 564MLA. Implementation of Multi-Head Latent Attention (MLA) mechanism. (By learning DeepSeek)
★ 8NVSHMEM-Tutorial. NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer
★ 195batch_invariant_ops. Python
★ 1.1kAttentionEngine. Python
★ 123nvshmem. NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
★ 567qwen600. Static suckless single batch CUDA-only qwen3-0.6B mini inference engine
★ 556SwiftTransformer. High performance Transformer implementation in C++.
★ 155swiftLLM. A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
★ 330Awesome-Efficient-MoE. Efficient Mixture of Experts for LLM Paper List
★ 184how-to-build-a-coding-agent. A workshop that teaches you how to build your own coding agent. Similar to Roo code, Cline, Amp, Cursor, Windsurf or OpenCode.
★ 5.8kEfficient_Attention_Survey. A Survey of Efficient Attention Methods: Hardware-efficient, Sparse, Compact, and Linear Attention
★ 305tilus. Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.
★ 490RLFromScratch. Python
★ 645cudaLLM. Python
★ 149flash-attention-with-sink. Python
★ 37nsa-release. An efficient implementation of the NSA (Native Sparse Attention) kernel
★ 134hilt. Python
★ 40NVSHEMEM. Sample Codes using NVSHMEM on Multi-GPU
★ 30grouped-query-attention-pytorch. (Unofficial) PyTorch implementation of grouped-query attention (GQA) from "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints" (https://arxiv.org/pdf/2305.13245.pdf)
★ 194qutlass. QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning
★ 195