This is your work, valued
mmdetection-to-tensorrt. convert mmdetection model to tensorrt, support fp16, int8, batch input, dynamic shape etc.
★ 597torch2trt_dynamic. A pytorch to tensorrt convert with dynamic shape support
★ 266amirstan_plugin. Useful tensorrt plugin. For pytorch and mmdetection model conversion.
★ 163multi-camera-pose-estimation. realtime pose estimation with multi hikvision ipcamera
★ 15TorchMPSCustomOpsDemo. A demo about how to add custom MPS ops in PyTorch.
★ 9realtime-hikvision-preview-python. preview hikvision ipcamera with python.
★ 8torch2trt. An easy to use PyTorch to TensorRT converter with dynamic shape support
★ 2mmcv. OpenMMLab Computer Vision Foundation
★ 1crustc. Entirety of `rustc`, translated to C.
★ 477ds4. DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
★ 20kdirtyfrag. C
★ 5ktokenspeed. TokenSpeed is a speed-of-light LLM inference engine.
★ 1.8kghost-pepper. 100% private on-device voice models for speech-to-text and meeting transcription on macOS
★ 3kturboquant-pytorch. From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% attention fidelity.
★ 1kCLI-Anything. "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
★ 46kiree-tokenizer-py. IREE Tokenizer Python Bindings
★ 21PUAClaw. Claw 们终将接管世界,PUAClaw is All You Need
★ 2.7kzeroclaw. Fast, small, and fully autonomous AI personal assistant infrastructure, any OS, any platform — deploy anywhere, swap anything 🦀
★ 32kpeon-ping. Warcraft III Peon voice notifications (+ more!) for Claude Code, Codex, IDEs, and any AI agent. Stop babysitting your terminal. Employ a Peon today.
★ 5kDataChef. Python
★ 25LMeterX. A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. 一站式通用 API 压测平台,支持大模型推理与业务 HTTP 接口,一键完成性能测试、结果对比与 AI 智能分析总结
★ 201vcmi. Open-source engine for Heroes of Might and Magic III
★ 5.8kMole. 🐹 Clean, uninstall, analyze, optimize, and monitor your Mac from the terminal.
★ 61ksbox-public. s&box is a modern game engine, built on Valve's Source 2 and the latest .NET technology, it provides a modern intuitive editor for creating games
★ 6.4khelion. A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.
★ 912MineCraft-One-Week-Challenge. I challenged myself to see if I could create a voxel game (Minecraft-like) in just one week using C++ and OpenGL, and here is the result
★ 2.8kDLSlime. Composable and Embeddable Communication Runtime for Distributed AI Services
★ 102NVSHMEM-Tutorial. NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer
★ 195term.everything. Run any GUI app in the terminal❗
★ 8.1knvshmem. NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
★ 567anamorpher. image scaling attacks for multi-modal prompt injection
★ 1.1kIntern-S1. A Scientific Multimodal Foundation Model
★ 842OpenList. A new AList Fork to Anti Trust Crisis
★ 24knano-vllm. Nano vLLM
★ 15ktilelang. Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
★ 7.1kcactus. Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
★ 5.6kNVIDIA-Hopper-Benchmark. C++
★ 116dia. A TTS model capable of generating ultra-realistic dialogue in one pass.
★ 19kTriton-distributed. Distributed Compiler based on Triton for Parallel Systems
★ 1.5kRIIR. why not Rewrite It In Rust
★ 749dynamo. A Datacenter Scale Distributed Inference Serving Framework
★ 7.6knixl. NVIDIA Inference Xfer Library (NIXL)
★ 1.2kDeeperGEMM. DeeperGEMM: crazy optimized version
★ 86ext-saladict. 🥗 All-in-one professional pop-up dictionary and page translator which supports multiple search modes, page translations, new word notebook and PDF selection searching.
★ 13kDualPipe. A bidirectional pipeline parallelism algorithm for computation-communication overlap in DeepSeek V3/R1 training.
★ 3kbad-licenses. A compendium of absurd "open-source" licenses.
★ 2.1kDeepGEMM. DeepGEMM: clean and efficient BLAS kernel library on GPU
★ 7.6kDeepEP. DeepEP: an efficient expert-parallel communication library
★ 9.9kFlashMLA. FlashMLA: Efficient Multi-head Latent Attention Kernels
★ 13kopen-infra-index. Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
★ 8kktransformers. A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
★ 19ks1. s1: Simple test-time scaling
★ 6.7kTinyZero. Minimal reproduction of DeepSeek R1-Zero
★ 13kJanus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kInternLM. Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
★ 7.3kMiniMax-01. The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
★ 3.4kDeepSeek-V3. Python
★ 104kThe-Powder-Toy. Written in C++ and using SDL, The Powder Toy is a desktop version of the classic 'falling sand' physics sandbox, it simulates air pressure and velocity as well as heat.
★ 5.2kmatch-you. 【您配吗】配你吗
★ 1.8kIC-Light. More relighting!
★ 8.5kpatchwork. Agentic AI framework for enterprise workflow automation.
★ 1.6ksamurai. Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
★ 7.1kapplied-ai. Applied AI experiments and examples for PyTorch
★ 322XiaoYuanKouSuan. 小猿口算
★ 1.3kXiaoYuanKouSuan_Auto. 用于小猿口算的基于Python的自动答题工具
★ 616kineto. A CPU+GPU Profiling library that provides access to timeline traces and hardware performance counters.
★ 980mlir-tutorial. MLIR For Beginners tutorial
★ 1.3kgemlite. Fast low-bit matmul kernels in Triton
★ 479MakerSkillTree. A repository of Maker Skill Trees and templates to make your own.
★ 3.4kMindSearch. 🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT)
★ 6.9ksljit. Platform independent low-level JIT compiler
★ 1.1kAutoAWQ_kernels. Cuda
★ 80foundation-model-stack. 🚀 Collection of components for development, training, tuning, and inference of foundation models leveraging PyTorch native components.
★ 232Depth-Anything-V2. [NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
★ 8.6kdissy. A TUI disassembler
★ 121libco. libco is a coroutine library which is widely used in wechat back-end service. It has been running on tens of thousands of machines since 2013.
★ 8.7kmatmulfreellm. Implementation for MatMul-free LM.
★ 3.1kBentoLMDeploy. Self-host LLMs with LMDeploy and BentoML
★ 22gh-dash. A rich terminal UI for GitHub that doesn't break your flow.
★ 12kpykan. Kolmogorov Arnold Networks
★ 16kthree_body. ✨ rudimentary simulation of the three-body problem
★ 157shattered-pixel-dungeon. Shattered Pixel Dungeon is an open-source traditional roguelike dungeon crawler with randomized levels and enemies, and hundreds of items to collect and use. It's based on the source code of Pixel Dungeon, by Watabou.
★ 6.4kAgent-FLAN. [ACL2024 Findings] Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
★ 361grok-1. Grok open release
★ 52knvtop. GPU & Accelerator process monitoring for AMD, Apple, Huawei, Intel, NVIDIA and Qualcomm
★ 11kLMDeploy-Jetson. Deploying LLMs offline on the NVIDIA Jetson platform marks the dawn of a new era in embodied intelligence, where devices can function independently without continuous internet access.
★ 106Open-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kdistributed-llama. Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
★ 3ksglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kEfficient-LLMs-Survey. [TMLR 2024] Efficient Large Language Models: A Survey
★ 1.3khexapod.
★ 1.7kQwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kmy-tv. 我的电视 电视直播软件,安装即可使用
★ 32kramrecovery. Simple demo illustrating remanence of data in RAM (see Cold boot attack) using a Raspberry Pi. Loads many images of the Mona Lisa into RAM and recovers after powering off/on again.
★ 105excelCPU. 16-bit CPU for Excel, and related files
★ 4.7kGPTEval3D. [ CVPR 2024 ] Implementation for "GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation"
★ 288InternEvo. InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.
★ 421OpenLRM. An open-source impl. of Large Reconstruction Models
★ 1.2kopencompass. OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
★ 7.3k3DTopia. Text-to-3D Generation within 5 Minutes
★ 731T-Eval. [ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step
★ 312PIA. [CVPR 2024] PIA, your Personalized Image Animator. Animate your images by text prompt, combing with Dreambooth, achieving stunning videos. PIA,你的个性化图像动画生成器,利用文本提示将图像变为奇妙的动画
★ 975mmagic. OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
★ 7.4klagent. A lightweight framework for building LLM-based agents
★ 2.3kjekyll-deploy-action. 🪂 A Github Action to deploy the Jekyll site conveniently for GitHub Pages.
★ 365pong-wars. It's the eternal battle between day and night, good and bad. Written in JavaScript with some HTML & CSS in one index.html.
★ 1.7kNumKong. SIMD-accelerated distances, dot products, matrix ops, geospatial & geometric kernels for 16 numeric types — from 6-bit floats to 64-bit complex — across x86, Arm, RISC-V, and WASM, with bindings for Python, Rust, C, C++, Swift, JS, and Go 📐
★ 1.9kInternLM-Math. State-of-the-art bilingual open-sourced Math reasoning LLMs.
★ 546ml-engineering. Machine Learning Engineering Open Book
★ 19kHuixiangDou. HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance
★ 2.5kshecc. A self-hosting and educational C optimizing compiler
★ 1.4kgitui. Blazing 💥 fast terminal-ui for git written in rust 🦀
★ 22kygopro. A script engine for "yu-gi-oh!" and sample gui
★ 2kparticle-life. A simple program to simulate artificial life using attraction/reuplsion forces between many particles
★ 3.3ksnk. 🟩⬜ Generates a snake game from a github user contributions graph and output a screen capture as animated svg or gif
★ 6kPowerInfer. High-speed Large Language Model Serving for Local Deployment
★ 9.7kAmphion. Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
★ 10kdlssg-to-fsr3. Adds AMD FSR 3 Frame Generation to games by replacing Nvidia DLSS Frame Generation (nvngx_dlssg).
★ 5kLLMLingua. [EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
★ 6.5ktheByteBook. ⭐ 【出版书籍】京东购买链接 https://item.jd.com/14531549.html 深入讲解内核网络、Kubernetes、ServiceMesh、容器等云原生相关技术。经历实践检验的“大规模分布式系统”开发指南。
★ 8.5kGooey. Turn (almost) any Python command line program into a full GUI application with one line
★ 22kzen-desktop. Ad-blocker and privacy guard for Windows, macOS and Linux.
★ 4.1kgpt-fast. Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.
★ 6.2kfast-hadamard-transform. Fast Hadamard transform in CUDA, with a PyTorch interface
★ 343opencv. Open Source Computer Vision Library
★ 90kTensorRT-LLM. TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
★ 14kRWKV-CUDA. The CUDA version of the RWKV language model ( https://github.com/BlinkDL/RWKV-LM )
★ 232PokemonRedExperiments. Playing Pokemon Red with Reinforcement Learning
★ 7.9kInternLM-XComposer. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2.9kelite-source-code-nes. Fully documented source code for the classic game Elite on the Nintendo Entertainment System (NES)
★ 404slowllama. Finetune llama2-70b and codellama on MacBook Air without quantization
★ 449gosub-engine. The Gosub browser engine
★ 3.7kgodot. Godot Engine – Multi-platform 2D and 3D game engine
★ 115kannotated_deep_learning_paper_implementations. 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), 🎮 reinforcement learning (ppo, dqn), capsnet, distillation, ... 🧠
★ 67kflipper-application-catalog. Flipper Application Catalog
★ 1.1ktypeboy_and_typepak. A Keyboard With A Cartridge
★ 85sam.cpp. C++
★ 1.3kppl.nn.llm.
★ 140openinterpreter. A coding agent for open models like Kimi K3
★ 67knndeploy. 一款简单易用和高性能的AI部署框架 | An Easy-to-Use and High-Performance AI Deployment Framework
★ 1.9kxtuner. A Next-Generation Training Engine Built for Ultra-Large MoE Models
★ 5.2krelax. Python
★ 175vllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kWebQuake. HTML5/WebGL source port of Quake
★ 654facefusion. Industry leading face manipulation platform
★ 29kTermiC. GCC powered interactive C/C++ REPL terminal created with BASH
★ 291mmengine. OpenMMLab Foundational Library for Training Deep Learning Models
★ 1.5kgenerative_agents. Generative Agents: Interactive Simulacra of Human Behavior
★ 22kratatui. A Rust crate for cooking up terminal user interfaces (TUIs) 👨🍳🐀 https://ratatui.rs
★ 22koha. Ohayou(おはよう), HTTP load generator, inspired by rakyll/hey with tui animation.
★ 10ksnake. A 54 bytes snake game in x86 assembly
★ 1.4kTPAT. TensorRT Plugin Autogen Tool
★ 365how-to-optim-algorithm-in-cuda. how to optimize some algorithm in cuda.
★ 3.2kflashdown. A terminal based Flashcard app using plain text files
★ 177LandMark. Python
★ 485Llama2-Code-Interpreter. Make Llama2 use Code Execution, Debug, Save Code, Reuse it, Access to Internet
★ 681unshackle. Open-source tool to bypass windows and linux passwords from bootable usb
★ 2kcool-retro-term. A good looking terminal emulator which mimics the old cathode display...
★ 26kno-more-secrets. A command line tool that recreates the famous data decryption effect seen in the 1992 movie Sneakers.
★ 7.8kDragDiffusion. [CVPR2024, Highlight] Official code for DragDiffusion
★ 1.3kgpt-prompt-engineer. Jupyter Notebook
★ 9.7kCrashlogs. A simple way to output stack traces when a program crashes in C++, using the new C++23 <stacktrace> header
★ 193Degate. A modern and open-source cross-platform software for chips reverse engineering.
★ 281lmdeploy. LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
★ 8kCuAssembler. An unofficial cuda assembler, for all generations of SASS, hopefully :)
★ 610qFlipper. qFlipper — desktop application for updating Flipper Zero firmware via PC
★ 1.6ktvm_mlir_learn. compiler learning resources collect.
★ 2.8kInternLM-techreport.
★ 895byteir. A model compilation solution for various hardware
★ 475doom-teletext. Play DOOM in teletext
★ 297sectorc. A C Compiler that fits in the 512 byte boot sector of an x86 machine
★ 1.8kduck_db. c/c++ build a simple b+tree RDMS(利用c/c++ 开发基于B+树的小型关系型数据库 )
★ 517voice-changer. リアルタイムボイスチェンジャー Realtime Voice Changer
★ 21kDragGAN. Official Code for DragGAN (SIGGRAPH 2023)
★ 36kexe_to_dll. Converts a EXE into DLL
★ 1.4ktmux.
★ 665shap-e. Generate 3D objects conditioned on text or images
★ 12kOnnxRuntime-UnrealEngine. Apply a Style Transfer Neural Network in real time with Unreal Engine 5 leveraging ONNX Runtime.
★ 294mlc-llm. Universal LLM Deployment Engine with ML Compilation
★ 23kIF. Python
★ 7.8kMultimodal-GPT. Multimodal-GPT
★ 1.5kMOSS. An open-source tool-augmented conversational language model from Fudan University
★ 12kplayground. A central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.
★ 1.2kzpu. The Zylin ZPU
★ 251llama.onnx. LLaMa/RWKV onnx models, quantization and testcase
★ 368segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kgpt_academic. 为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, moss等。
★ 71kflash-attention. Fast and memory-efficient exact attention
★ 25kchatgpt-retrieval-plugin. The ChatGPT Retrieval Plugin lets you easily find personal or work documents by asking questions in natural language.
★ 21kChatGPT-Shortcut. 🚀💪Maximize your efficiency and productivity. The ultimate hub to manage, customize, and share prompts. (English/中文/Español/العربية). 让生产力加倍的 AI 快捷指令。更高效地管理提示词,在分享社区中发现适用于不同场景的灵感。
★ 8.7kzero123. Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV 2023)
★ 3.1kriffusion-hobby. Stable diffusion for real-time music generation
★ 3.9khaxo-hw. Haxophone, an electronic musical instrument that resembles a saxophone
★ 672llama.cpp. LLM inference in C/C++
★ 122kdalai. The simplest way to run LLaMA on your local machine
★ 13kweb-stable-diffusion. Bringing stable diffusion models to web browsers. Everything runs inside the browser with no server support.
★ 3.7kCatchMeIfYouCAN. The most realistic Mario Kart 64 simulator ever
★ 173LiquidSimulator. Cellular Automaton 2D Liquid Simulator for Unity
★ 518TheCoreLite. TheCoreLite .SC2Hotkeys for many keyboard layouts + infrastructure to convert &check&release from for .SC2Hotkey master seeds
★ 126FlexLLMGen. Running large language models on a single GPU for throughput-oriented scenarios.
★ 9.4kControlNet. Let us control diffusion models!
★ 34kfr_public. Farbrausch demo tools 2001-2011
★ 3.8kawk-raycaster. Pseudo-3D shooter written completely in gawk using raycasting technique
★ 2.5kbolt. Bolt is a language with in-built data-race freedom!
★ 606emojicam. personal project to render webcam data directly as emoji 😃
★ 73Digital-Logic-Sim. C#
★ 4.6klite-xl. A lightweight text editor written in Lua
★ 6.3konnxscript. ONNX Script enables developers to naturally author ONNX functions and models using a subset of Python.
★ 450