This is your work, valued
ML inference performance: compilers, runtimes, and GPU kernels
lintel. A Python module to decode video frames directly, using the FFmpeg C API.
★ 260LOHO. Demo code for "LOHO: Latent Optimization of Hairstyles via Orthogonalization".
★ 171SSTVOS. Training code for "SSTVOS: Sparse Spatiotemporal Transformers for Video Object Segmentation"
★ 88ENAS-pytorch. PyTorch implementation of "Efficient Neural Architecture Search via Parameters Sharing"
★ 6mobile-pruning. Python
★ 6ml-reading-notes-improved-system. A set of reading summaries related to machine learning.
★ 4nerfies-to-3d-generative. Python
★ 4master-thesis-fusion. TeX
★ 4arm-disassembler. Disassembler for ARMv5 architecture
★ 3personal-website. HTML
★ 3MichiGAN. MichiGAN: Multi-Input-Conditioned Hair Image Generation for Portrait Editing (SIGGRAPH 2020)
★ 2imgaug. Image augmentation for machine learning experiments.
★ 2ridesharing-taxicab-scheduler. C++
★ 2learning-sql. Originally from https://resources.oreilly.com/examples/9780596007270/
★ 2advent-of-code.
★ 1vscode-debug-mixed-python-cpp. Based on: https://nadiah.org/2020/03/01/example-debug-mixed-python-c-in-visual-studio-code/
★ 1cp-algo. C++
★ 1mlir-www. SCSS
★ 1c-extension-tutorial. How to Write and Debug C Extension Modules
★ 1mlir-hacking. MLIR
★ 1python-c-extension-hacking. C
★ 1pybind11-hacking.
★ 1tensorflow-special-octo-spoon. Python
★ 1go.dev-tutorials. Go
★ 1sparse-spatiotemporal-transformer.
★ 1ml-programming-problems. Python
★ 1diffusercam. TeX
★ 1learn-reactjs. JavaScript
★ 1quartus-ii-projects. Verilog
★ 1concurrency-in-action. C++
★ 1random-vimrc-etc. Vim Script
★ 1python-dabbling. I dabble
★ 1phd-thesis. HTML
★ 1mit.6824.
★ 1bassoon. A parameter server compatible with PyTorch optimizers.
★ 1web-dev-zero-to-hero. HTML
★ 1x86_Towers_of_Hanoi. Assembly
★ 1verilog-hdl-palnitkar. Verilog
★ 1apc-literate-chainsaw. Python
★ 1AST_PrettyPrinter. Files to pretty-print A1.java, A2.java and A3.java from CS2S03 Assignment 1, using an AST.
★ 1face-action-unit-detection. Python
★ 1top-work. https://dukebw.github.io/top-work/
★ 1Connect_4. An electronic version of the popular two-player connection game
★ 1programming-challenges-skiena. C
★ 1stern. ⎈ Multi pod and container log tailing for Kubernetes -- Friendly fork of https://github.com/wercker/stern
★ 4.8kFlyDSL. FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.
★ 258harnessgym. Iterative agent harness improvement: run a coding agent on a hard task, generate the reusable tooling it was missing, qualify it, and replay fresh sessions with it activated. Works with Codex and Claude Code.
★ 40aiperf. AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
★ 492DeepSpec. DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
★ 6.8kSana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6kgit-pile. Stacked diff support for GitHub workflows
★ 181MSA. Python
★ 392SGLang-FluentLLM. Python
★ 114turboquant. TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
★ 1.7kOpenCode-goal-plugin. Durable, guarded goal workflows for OpenCode with persistence, safety limits, agent tools, and evidence-gated completion.
★ 216ideogram4. Ideogram 4: Open image model at the forefront of design
★ 2.6kFastGen. NVIDIA FastGen: Fast Generation from Diffusion Models
★ 914export-python. Conveniently export torch.compile compiled products into self-contained Python files
★ 35tokenspeed. TokenSpeed is a speed-of-light LLM inference engine.
★ 1.8kml-cookbook. Ready-to-use ML training recipes to help you build and deploy models on Baseten.
★ 60Stirrup. The lightweight framework for building agents
★ 529aiter. AI Tensor Engine for ROCm
★ 510eza. A modern alternative to ls
★ 23kpyrefly. A fast type checker and language server for Python
★ 6.8kCobraML2. Performant kernels, and other ML Systems integrations
★ 6triton-tutorial. From a+b to sparsemax(QK^T)V in Triton!
★ 34original_performance_takehome. Anthropic's original performance take-home, now open for you to try!
★ 4.1ktilelang. Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
★ 7.1kTileKernels. A kernel library written in tilelang
★ 1.7kcuLA. CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.
★ 535codex-plugin-cc. Use Codex from Claude Code to review code or delegate tasks.
★ 31kintra-kernel-profiler. Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.
★ 123InferenceX-app. Dashboard for InferenceX™, Open Source Continuous Inference
★ 37steplaw. Python
★ 235skills. Skills for Real Engineers. Straight from my .agents directory.
★ 198kSpecForge. Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
★ 1klat.md. Agent Lattice: a knowledge graph for your codebase, written in markdown.
★ 1.8kbtop. A monitor of resources
★ 34kcolloquium. A markdown native slides tool for academics building with agents.
★ 236rtk. CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
★ 74kopencode-gemini-auth. Gemini auth plugin for opencode
★ 1.7kTorchSpec. A PyTorch native library for training speculative decoding models
★ 217pi. AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
★ 82kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93kcli. Google Workspace CLI — one command-line tool for Drive, Gmail, Calendar, Sheets, Docs, Chat, Admin, and more. Dynamically built from Google Discovery Service. Includes AI agent skills.
★ 30kInferenceX. Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
★ 1.3kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kocto.nvim. Edit and review GitHub issues and pull requests from the comfort of your favorite editor
★ 3.3kKimi-Vendor-Verifier. Kimi-Vendor-Verifier
★ 95excalidraw-mcp. Fast and streamable Excalidraw MCP App
★ 5koh-my-openagent. omo/lazycodex: The coding agent for tokenmaxxers;the one and only agent harness for complex codebases. For your Codex, for your OpenCode
★ 67kvllm-omni. A framework for efficient model inference with omni-modality models
★ 5.8kquill. Asynchronous Low Latency C++ Logging Library
★ 3kremote-nvim.nvim. Remote development in Neovim 🔥
★ 1.3kzellij-nav.nvim. Seamless navigation between Neovim windows and Zellij panes.
★ 233mutagen. Fast file synchronization and network forwarding for remote development
★ 4.3kopentui. OpenTUI is a library for building terminal user interfaces (TUIs)
★ 13kDeepSeek-OCR-2. Visual Causal Flow
★ 3.2kkimi-cli. Kimi Code CLI is your next CLI agent.
★ 11ktoad. A unified interface for AI in your terminal.
★ 3.3krich. Rich is a Python library for rich text and beautiful formatting in the terminal.
★ 57ktextual. The lean application framework for Python. Build sophisticated user interfaces with a simple Python API. Run your apps in the terminal and a web browser.
★ 37kvibetensor. Our first fully AI generated deep learning system
★ 635kvpress. LLM KV cache compression made easy
★ 1.2kmirage. Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
★ 2.4kqmd. mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local
★ 28kLongBench. LongBench v2 and LongBench (ACL 25'&24')
★ 1.2kRULER. This repo contains the source code for RULER: What’s the Real Context Size of Your Long-Context Language Models?
★ 1.6kGLM-ASR-WebUI. A Web UI for easy subtitle using glm asr model.
★ 7Step-Audio2. Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
★ 1.5kSenseVoice. Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
★ 9kGLM-ASR. GLM-ASR-Nano: A robust, open-source speech recognition model with 1.5B parameters
★ 837Fun-ASR. Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
★ 1.5klightning-whisper-mlx. An extremely fast implementation of whisper optimized for Apple Silicon using MLX.
★ 951argmax-oss-swift. On-device Speech AI for Apple Silicon
★ 6.3kyoutube-transcript-api. This is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it does not require an API key nor a headless browser, like other selenium based solutions do!
★ 8kfaster-whisper. Faster Whisper transcription with CTranslate2
★ 25kRocketKV. [ICML 2025] RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
★ 50nmoe. MoE training for Me and You and maybe other people
★ 396mini-sglang. A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
★ 4.7knanotrace. Low overhead tracing library and trace visualizer for pipelined CUDA kernels
★ 135iris. AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming
★ 194tvm-ffi. Open ABI and FFI for Machine Learning Systems
★ 439Block-Sparse-Flash-Attention. C++
★ 34quack. A Quirky Assortment of CuTe Kernels
★ 1.1kMegakernels. Kernels, of the mega variety :)
★ 788gpu-experiments. A collection of GPU experiments and benchmarks for my personal understanding and research.
★ 36Alpha-MoE. Cuda
★ 74comm_scope. NUMA-aware multi-CPU multi-GPU data transfer benchmarks
★ 28pyATF. Python
★ 19slime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kkineto. A CPU+GPU Profiling library that provides access to timeline traces and hardware performance counters.
★ 980LPLB. An early research stage expert-parallel load balancer for MoE models based on linear programming.
★ 522grok-cli. An open-source coding agent for the Grok API
★ 3.4kmori. Modular RDMA Interface
★ 164HipKittens. Fast and Furious AMD Kernels
★ 447folly. An open-source C++ library developed and used at Facebook.
★ 30kllm-inference-handbook. Everything you need to know about LLM inference
★ 367jsonargparse. Minimal effort CLIs derived from type hints and parse from command line, config files and environment variables
★ 428FireRedASR. Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
★ 2kxllm. A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
★ 1.5kai-performance-engineering. Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.
★ 1.8kNVSentinel. NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated computing environments
★ 358diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34khorovod. Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.
★ 15klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kwerkzeug. The comprehensive WSGI web application library.
★ 6.9ksj.h. A tiny little JSON parsing library
★ 1.5kcline. Autonomous coding agent as an SDK, IDE extension, or CLI assistant.
★ 65kSysinternals-jcd. Sysinternals jcd (jump change directory) is a Rust-based command-line tool that provides enhanced directory navigation with substring matching and smart selection. It's like the cd command, but with superpowers!
★ 173yamoe. 🔀 yet another mixture of experts
★ 23cyclopts. Intuitive, easy CLIs based on python type hints.
★ 1.2knvshmem. NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
★ 567qutlass. QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning
★ 195guidellm. Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
★ 1.5kdynolog. Dynolog is a telemetry daemon for performance monitoring and tracing. It exports metrics from different components in the system like the linux kernel, CPU, disks, Intel PT, GPUs etc. Dynolog also integrates with pytorch and can trigger traces for distributed training applications.
★ 376Flash-Sparse-Attention. 🚀🚀 Efficient implementations of Native Sparse Attention
★ 622markitdown. Python tool for converting files and office documents to Markdown.
★ 171kclaude-squad. Manage multiple AI terminal agents like Claude Code, Codex, OpenCode, and Amp.
★ 8.2kaiconfigurator. Offline optimization of your disaggregated Dynamo graph
★ 384vllm-cli. A command-line interface tool for serving LLM using vLLM.
★ 507marin. Open-source framework for the research and development of foundation models.
★ 1.2ktorch-mojo-backend. Mojo
★ 24mojo-gpu-puzzles. Learn GPU Programming in Mojo🔥 by Solving Puzzles
★ 3560xtools. 0x.Tools: X-Ray vision for Linux systems
★ 1.8kNSA-Test. NSA Triton Kernels written with GPT5 and Opus 4.1
★ 70StepMesh. C++
★ 381teamtype. Peer-to-peer, editor-agnostic collaborative editing of local text files.
★ 1.9kcua. Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
★ 21kLiveCodeBench-Pro. Python
★ 176asciinema. Terminal session recorder, streamer and player 📹
★ 18kqwen-code. An open-source AI coding agent that lives in your terminal.
★ 26kharmony. Renderer for the harmony response format to be used with gpt-oss
★ 4.5kGLM-4. GLM-4 series: Open Multilingual Multimodal Chat LMs | 开源多语言多模态对话模型
★ 7.1kkubectx. Faster way to switch between clusters and namespaces in kubectl
★ 20kopencode. The open source coding agent.
★ 192kDeepEP_ibrc_dual-ports_multiQP. Aims to implement dual-port and multi-qp solutions in deepEP ibrc transport
★ 75delayed-streams-modeling. Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.
★ 3kui-terminal-mojo. terminal ui mojo programming language
★ 10live-stats-for-linux-based-operating-systems-in-mojo. live stats dashboad for linux based operating systems in mojo
★ 4xyne. AI-first Search & Answer Engine for work. Open-source alternative to Glean.
★ 674uccl. UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
★ 1.5kPDF-Extract-Kit. A Comprehensive Toolkit for High-Quality PDF Content Extraction
★ 9.9kPaddle. PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
★ 24kERNIE. The official repository for ERNIE 4.5 and ERNIEKit – its industrial-grade development toolkit based on PaddlePaddle.
★ 7.7kyet-another-bench-script. YABS - a simple bash script to estimate Linux server performance using fio, iperf3, & Geekbench
★ 6.6kmodular-hack-weekend-2506. My work for Modular Hack Weekend, June 2025
★ 2exo2-artifact. Artifact Evaluation for ASPLOS '25 Exo 2 paper
★ 5ish. Alignment-based filtering CLI tool
★ 62rocSHMEM. [DEPRECATED] Moved to ROCm/rocm-systems repo
★ 145ome. Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
★ 482gemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kclaude-code.nvim. Seamless integration between Claude Code AI assistant and Neovim
★ 2.1knano-vllm. Nano vLLM
★ 15kanki-mcp-server. An MCP server for Anki
★ 184anki-mcp-server. A model context protocol server that connects to Anki through AnkiConnect
★ 83anki-mcp-server. MCP server for Anki via AnkiConnect
★ 246niri. A scrollable-tiling Wayland compositor.
★ 27kSWE-agent. SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]
★ 20kmlx-lm. Run LLMs with MLX
★ 6.5kInternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kMNN. MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
★ 16kESFT. Expert Specialized Fine-Tuning
★ 740peft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kMulti-MoE. Python
★ 1llm-d. Achieve state of the art inference performance with modern accelerators on Kubernetes
★ 4kNuMojo. NuMojo is a library for numerical computing in Mojo 🔥 similar to numpy in Python.
★ 222skypilot. The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
★ 10kmodular-hacks-physics-sim. implementation of physics sim using mojo
★ 5monolithic-sup. Mojo
★ 3capnproto. Cap'n Proto serialization/RPC system - core tools and C++ library
★ 13kliburing. Library providing helpers for the Linux kernel io_uring support
★ 3.7kgdrcopy. A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology
★ 1.4kucx. Unified Communication X (mailing list - https://elist.ornl.gov/mailman/listinfo/ucx-group)
★ 1.7kcompressed-tensors. A safetensors extension to efficiently store sparse quantized tensors on disk
★ 308nixl. NVIDIA Inference Xfer Library (NIXL)
★ 1.2kmermaid-cli. Command line tool for the Mermaid library
★ 4.9kKVDirect. Code for KVDirect: Distributed Disaggregated LLM Inference
★ 10dust. A more intuitive version of du in rust
★ 12kBitNet. Official inference framework for 1-bit LLMs
★ 40kmojo-operation-template. A starter template for working on new Mojo CPU / GPU operations
★ 1Mooncake. Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6.1kentropix. Entropy Based Sampling and Parallel CoT Decoding
★ 3.4kCMake. Mirror of CMake upstream repository
★ 8kDistServe. Disaggregated serving system for Large Language Models (LLMs).
★ 826splitwise-sim. LLM serving cluster simulator
★ 157codex. Lightweight coding agent that runs in your terminal
★ 103kprofile-data. Analyze computation-communication overlap in V3/R1.
★ 1.2kmax-mamba. port of mamba architecture to MAX+MOJO
★ 2pplx-kernels. Perplexity GPU Kernels
★ 597vocos. Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
★ 1.1kDeepSeek-MoE. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
★ 2kKAI-Scheduler. KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale
★ 1.4ktensorizer. Module, Model, and Tensor Serialization/Deserialization
★ 319fastsafetensors. High-performance safetensors model loader
★ 162InfiniStore. KV cache store for distributed LLM inference
★ 424train-tk. train with kittens!
★ 67v6d. vineyard (v6d): an in-memory immutable data manager. (Project under CNCF, TAG-Storage)
★ 961text-generation-inference. Large Language Model Text Generation Inference
★ 11kSpatialLM. [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
★ 4.7kBayLing-Speech. LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
★ 3.1kIsaac-GR00T. NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7.7kdynamo. A Datacenter Scale Distributed Inference Serving Framework
★ 7.7kglake. GLake: optimizing GPU memory management and IO transmission.
★ 501