This is your work, valued
yolort. yolort is a runtime stack for yolov5 on specialized accelerators such as tensorrt, libtorch, onnxruntime, tvm and ncnn.
★ 729sightseq. Computer vision tools for fairseq, containing PyTorch implementation of text recognition and object detection
★ 123demonet. Yet another ssd, with its runtime stack for libtorch, onnx and specialized accelerators.
★ 26shufaCV. Python
★ 26simple-faster-rcnn. Object detection from torchvision, just make it more convenient to do some experiments.
★ 8huo. 🔥 日出江花红胜火,春来江水绿如蓝
★ 5zhiqwang.github.io. Codes and Notes
★ 5fairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
★ 5yir. yir
★ 4yolov5. YOLOv5 in PyTorch > ONNX > CoreML > TFLite
★ 3pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
★ 2zhiqwang. 👋 🌎
★ 2nanodet. ⚡Super fast and lightweight anchor-free object detection model. 🔥Only 1.8MB and run 97FPS on cellphone🔥
★ 2vision. Datasets, Transforms and Models specific to Computer Vision
★ 2keras2ncnn. A keras h5df to ncnn model converter
★ 1TensorRT. TensorRT is a C++ library for high performance inference on NVIDIA GPUs and deep learning accelerators.
★ 1sightsee. Sightsee is a thin CLI wrapper around netron for opening model files in a browser.
★ 1detr. End-to-End Object Detection with Transformers
★ 1torchscript-example. Example CMake project for TorchScript
★ 1ncnn. ncnn is a high-performance neural network inference framework optimized for the mobile platform
★ 1coros-mcp. MCP server for AI assistants to read and manage Coros fitness data: sleep, HRV, daily metrics, activities, and structured workouts via the unofficial Coros API
★ 104cutile-examples. cutile kernel examples
★ 51Happycapy-skills. A curated collection of high-quality Claude Code skills to enhance your development workflow
★ 138edict. 🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
★ 16kjax-js. JAX in JavaScript – ML library for the web, running on WebGPU & Wasm
★ 891LeetCUDA. LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
★ 12kSTAR. STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling Models
★ 52tilelang-puzzles. Learning TileLang with 10 puzzles!
★ 354onnxscript. ONNX Script enables developers to naturally author ONNX functions and models using a subset of Python.
★ 450DLCompiler. triton for dsa
★ 68tritonbench. Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.
★ 36310peaks-around-Beijing-ranking. 徒步强国京畿十峰挑战赛 排名可视化
★ 1ir-py. Efficient in-memory representation for ONNX, in Python
★ 45aurora. Implementation of the Aurora model for Earth system forecasting
★ 978tilelang. Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
★ 7.1kCUDA-GEMM-Optimization. CUDA Matrix Multiplication Optimization
★ 278HumanTOMATO. [ICML 2024] 🍅HumanTOMATO: Text-aligned Whole-body Motion Generation
★ 362OnnxSlim. A Toolkit to Help Optimize Large Onnx Model
★ 166nndeploy. 一款简单易用和高性能的AI部署框架 | An Easy-to-Use and High-Performance AI Deployment Framework
★ 1.9kpycorrector. pycorrector is a toolkit for text error correction. 文本纠错,实现了Kenlm,T5,MacBERT,ChatGLM3,Qwen2.5等模型应用在纠错场景,开箱即用。
★ 6.5kTinyGPT. Tiny C++ LLM inference implementation from scratch
★ 121CTranslate2. Fast inference engine for Transformer models
★ 4.6kCT2Hair. This is the official implementation of CT2Hair High-fidelity 3D Hair Modeling Using Computed Tomography.
★ 221jaxonnxruntime. A user-friendly tool chain that enables the seamless execution of ONNX models using JAX as the backend.
★ 136LightLLM. LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
★ 4.2klibflash_attn. C++
★ 14llama2.c. Inference Llama 2 in one file of pure C
★ 20kllama. Inference code for Llama models
★ 60klovely-tensors. Tensors, for human consumption
★ 1.4kco-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5kutilsd. Common deep learning utils.
★ 18open-resume. OpenResume is a powerful open-source resume builder and resume parser. https://open-resume.com/
★ 8.8kYOLOX. YOLOX is a high-performance anchor-free YOLO, exceeding yolov3~v5 with MegEngine, ONNX, TensorRT, ncnn, and OpenVINO supported. Documentation: https://yolox.readthedocs.io/
★ 11khidet. An open-source efficient deep learning framework/compiler, written in python.
★ 743Lidar_AI_Solution. A project demonstrating Lidar related AI solutions, including three GPU accelerated Lidar/camera DL networks (PointPillars, CenterPoint, BEVFusion) and the related libs (cuPCL, 3D SparseConvolution, YUV2RGB, cuOSD,).
★ 1.9kpytorch-pfn-extras. Supplementary components to accelerate research and development in PyTorch
★ 283pyinstrument. 🚴 Call stack profiler for Python. Shows you why your code is slow!
★ 8kNART. NART = NART is not A RunTime, a deep learning inference framework.
★ 37Dipoorlet. Offline Quantization Tools for Deploy.
★ 143onnx-extended. New operators for the ReferenceEvaluator, new kernels for onnxruntime, CPU, CUDA
★ 36Grounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55ksympy. A computer algebra system written in pure Python
★ 15kpytorch-lightning. Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
★ 31kPaDiff. Paddle Automatically Diff Precision Toolkits.
★ 53FlexLLMGen. Running large language models on a single GPU for throughput-oriented scenarios.
★ 9.4kFlyCV. FlyCV is a high-performance library for processing computer visual tasks.
★ 596ultralytics. Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
★ 60kcmake-examples. Useful CMake Examples
★ 13ktorchview. torchview: visualize pytorch models
★ 1.1kneural-compressor. SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
★ 2.7kRAM-multiprocess-dataloader. Demystify RAM Usage in Multi-Process Data Loaders
★ 206cpython. The Python programming language
★ 74kchaiNNer. A node-based image processing GUI aimed at making chaining image processing tasks easy and customizable. Born as an AI upscaling application, chaiNNer has grown into an extremely flexible and powerful programmatic image processing application.
★ 5.9kignite. High-level library to help with training and evaluating neural networks in PyTorch flexibly and transparently.
★ 4.8koptimum-intel. 🤗 Optimum Intel: Accelerate inference with Intel optimization tools
★ 609autocut. 用文本编辑器剪视频
★ 7.8kMegCC. MegCC是一个运行时超轻量,高效,移植简单的深度学习模型编译器
★ 483ax-pipeline. The Pipeline example based on AX650N/AX8850 shows the software development skills of Image Processing, NPU, Codec, and Display modules, which is helpful for users to develop their own multimedia applications.
★ 20FastDeploy. High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7kToMe. A method to increase the speed and lower the memory footprint of existing vision transformers.
★ 1.2kkid-chinese-stories. HTML
★ 3LM-Kernel-FT. A Kernel-Based View of Language Model Fine-Tuning https://arxiv.org/abs/2210.05643
★ 78torch-mechanics. Amateur experiments with autodiff mechanics simulators
★ 8fiddle. Python
★ 386PMN. [TPAMI 2023 / ACMMM 2022 Best Paper Runner-Up Award] Learnability Enhancement for Low-light Raw Denoising: Where Paired Real Data Meets Noise Modeling (a Data Perspective)
★ 180whisper.cpp. Port of OpenAI's Whisper model in C/C++
★ 52kAITemplate. AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.
★ 4.7kCuAssembler. An unofficial cuda assembler, for all generations of SASS, hopefully :)
★ 85sherpa. Speech-to-text server framework with next-gen Kaldi
★ 965detrex. detrex is a research platform for DETR-based object detection, segmentation, pose estimation and other visual recognition tasks.
★ 2.3kshumai. Fast Differentiable Tensor Library in JavaScript and TypeScript with Bun + Flashlight
★ 1.2kmmengine. OpenMMLab Foundational Library for Training Deep Learning Models
★ 1.5kkeras-cv. Industry-strength Computer Vision workflows with Keras
★ 1.1kSparsebit. A model compression and acceleration toolbox based on pytorch.
★ 331XMem. [ECCV 2022] XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model
★ 2kyolov5-rt-stack. (中文)yolort is a runtime stack for yolov5 on specialized accelerators such as libtorch, onnxruntime, tensorrt, tvm and ncnn.
★ 1onnx2torch. Convert ONNX models to PyTorch.
★ 738numpy. The fundamental package for scientific computing with Python.
★ 32konnx-modifier. A tool to modify ONNX models in a visualization fashion, based on Netron and Flask.
★ 1.6kJNeRF. JNeRF is a NeRF benchmark based on Jittor. JNeRF re-implemented instant-ngp and achieved same performance with original paper.
★ 642flash-attention. Fast and memory-efficient exact attention
★ 25kMegPeak. C++
★ 256symforce. Fast symbolic computation, code generation, and nonlinear optimization for robotics
★ 1.6kFasterTransformer. Transformer related optimization, including BERT, GPT
★ 6.4kax-samples. Samples code for world class Artificial Intelligence SoCs for computer vision applications.
★ 297PiPPy. Pipeline Parallelism for PyTorch
★ 786Vitis-AI. Vitis AI is Xilinx’s development stack for AI inference on Xilinx hardware platforms, including both edge devices and Alveo cards.
★ 1.8kopen-gpu-kernel-modules. NVIDIA Linux open GPU kernel module source
★ 17kLightGBM. A fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.
★ 19kAI-System. System for AI Education Resource.
★ 4.3kmetaseq. Repo for external large-scale work
★ 6.6kTensor-Puzzles. Solve puzzles. Improve your pytorch.
★ 4.3kNAFNet. The state-of-the-art image restoration model without nonlinear activation functions.
★ 3.1ksimple-onnx-processing-tools. A set of simple tools for splitting, merging, OP deletion, size compression, rewriting attributes and constants, OP generation, change opset, change to the specified input order, addition of OP, RGB to BGR conversion, change batch size, batch rename of OP, and JSON convertion for ONNX models.
★ 305PINTO_model_zoo. A repository for storing models that have been inter-converted between various frameworks. Supported frameworks are TensorFlow, PyTorch, ONNX, OpenVINO, TFJS, TFTRT, TensorFlowLite (Float32/16/INT8), EdgeTPU, CoreML.
★ 4.6ksog4onnx. Simple ONNX operation generator. Simple Operation Generator for ONNX.
★ 7torchopt. TorchOpt is an efficient library for differentiable optimization built upon PyTorch.
★ 635nanobind. nanobind: tiny and efficient C++/Python bindings
★ 3.6ktorchrec. Pytorch domain library for recommendation systems
★ 2.6kmkposters. Make posters from Markdown files.
★ 379mmrotate. OpenMMLab Rotated Object Detection Toolbox and Benchmark
★ 2.2kMatX. An efficient C++20 GPU numerical computing library with Python-like syntax
★ 1.4kkapao. KAPAO is an efficient single-stage human pose estimation model that detects keypoints and poses as objects and fuses the detections to predict human poses.
★ 770ncnn-small-board. ncnn benchmark on various single board computers
★ 166lite.ai.toolkit. A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.
★ 4.4krecipes. Recipes are a standard, well supported set of blueprints for machine learning engineers to rapidly train models using the latest research techniques without significant engineering overhead.Specifically, recipes aims to provide- Consistent access to pre-trained SOTA models ready for production- Reference implementations for SOTA research reproducibility, and infrastructure to guarantee correctness, efficiency, and interoperability.
★ 351mmdeploy. OpenMMLab Model Deployment Framework
★ 3.1knndet2. Python
★ 38instant-ngp. Instant neural graphics primitives: lightning fast NeRF and more
★ 18ktvm-cutlass-eval. Python
★ 41mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4kncnn-editor. ncnn和pnnx格式编辑器
★ 146amgcl. C++ library for solving large sparse linear systems with algebraic multigrid method
★ 872pyamg. Algebraic Multigrid Solvers in Python
★ 653ppq. PPL Quantization Tool (PPQ) is a powerful offline neural network quantization tool.
★ 1.8ktree-math. Mathematical operations for JAX pytrees
★ 210svox2. Plenoxels: Radiance Fields without Neural Networks
★ 2.9kBladeDISC. BladeDISC is an end-to-end DynamIc Shape Compiler project for machine learning workloads.
★ 931glide-text2im. GLIDE: a diffusion-based text-conditional image synthesis model
★ 3.7knipy. Neuroimaging in Python FMRI analysis package
★ 414vision. Datasets, Transforms and Models specific to Computer Vision
★ 18kyolov5. YOLOv5 in PyTorch > ONNX > CoreML > iOS
★ 220deepstream-video-pipeline. Python
★ 113tiny-cuda-nn. Lightning fast C++/CUDA neural network framework
★ 4.5kpython-lib-stats. Find usage statistics (imports, function calls, attribute access) for Python code-bases
★ 14torchscript-cmake-example. Example CMake project for TorchScript
★ 2SoftTeacher. Semi-Supervised Learning, Object Detection, ICCV2021
★ 926RAFT. Python
★ 4.1kpython-patterns. A collection of design patterns/idioms in Python
★ 43kunilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kMNN. MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
★ 16ktract. Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference
★ 3kYOLOv5_NCNN. 🍅 Deploy ncnn on mobile phones. Support Android and iOS. 移动端ncnn部署,支持Android与iOS。
★ 1.6kTinyNeuralNetwork. TinyNeuralNetwork is an efficient and easy-to-use deep learning model compression framework.
★ 880cvat. Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaboration, analytics, and developer APIs.
★ 16ktoroidal. a lightweight transformer library for PyTorch
★ 71aimet. AIMET is a library that provides advanced quantization and compression techniques for trained neural network models.
★ 2.7kpytorch-struct. Fast, general, and tested differentiable structured prediction in PyTorch
★ 1.1kghstack. Submit stacked diffs to GitHub on the command line
★ 1kaction-protobuf. CMake
★ 3xformers. Hackable and optimized Transformers building blocks, supporting a composable construction.
★ 11kmatplotlib. matplotlib: plotting with Python
★ 23kimnodes. A small, dependency-free node editor for dear imgui
★ 2.5ktaichi_tetris. Python
★ 58shufaCV. Python
★ 26calyx. Intermediate Language (IL) for Hardware Accelerator Generators
★ 608torchdynamo. A Python-level JIT compiler designed to make unmodified PyTorch programs faster.
★ 1.1kmujoco. Multi-Joint dynamics with Contact. A general purpose physics simulator.
★ 14kTorchSharp. A .NET library that provides access to the library that powers PyTorch.
★ 1.8knogil. Multithreaded Python without the GIL
★ 2.9kPyAV. Pythonic bindings for FFmpeg's libraries.
★ 3.3kdapi-model-versioning. RFC for Model Versioning across all PyTorch Domain libraries
★ 2scikit-learn. scikit-learn: machine learning in Python
★ 67kyolov5. YOLOv5 in PyTorch > ONNX > CoreML > TFLite
★ 3data. A PyTorch repo for data loading and utilities to be shared by the PyTorch domain libraries.
★ 1.3ktf2-detection-to-tvm. Python
★ 6buddy-mlir. An MLIR-based compiler framework bridges DSLs (domain-specific languages) to DSAs (domain-specific architectures).
★ 746yolov5-onnxruntime. YOLOv5 ONNX Runtime C++ inference code.
★ 288keras2ncnn. A keras h5df to ncnn model converter
★ 88PBRVulkan. Vulkan Real-time Path Tracer Engine
★ 527Kalman-and-Bayesian-Filters-in-Python. Kalman Filter book using Jupyter Notebook. Focuses on building intuition and experience, not formal proofs. Includes Kalman filters,extended Kalman filters, unscented Kalman filters, particle filters, and more. All exercises include solutions.
★ 19ktk-fangsong-font. 剔骨仿宋: Experimental Fang Song style Chinese font
★ 124bolt. 10x faster matrix and vector operations
★ 2.5klibigl. Simple MPL-2.0-licensed C++ geometry processing library.
★ 5.1kremake. やり直すんだ。そして、次はうまくやる。
★ 10kruntime. A performant and modular runtime for TensorFlow
★ 753mlir-aie. A close-to-metal Python API for programming AMD Ryzen™ AI NPUs (AI Engines), built on an open-source MLIR-based compiler toolchain.
★ 670pyxir. Python
★ 37pillow-simd. The friendly PIL fork
★ 2.3kpytorch-tools. Tool box for PyTorch
★ 187NNPACK. Acceleration package for neural networks on multi-core CPUs
★ 1.7kncnn. ncnn is a high-performance neural network inference framework optimized for the mobile platform
★ 126physics1. TeX
★ 298pnnx. PyTorch Neural Network eXchange
★ 709nni. An open source AutoML toolkit for automate machine learning lifecycle, including feature engineering, neural architecture search, model compression and hyper-parameter tuning.
★ 14knn-Meter. A DNN inference latency prediction toolkit for accurately modeling and predicting the latency on diverse edge devices.
★ 364xla. Enabling PyTorch on XLA Devices (e.g. Google TPU)
★ 2.8kpytorchic-bert. Pytorch Implementation of Google BERT
★ 600hora. 🚀 efficient approximate nearest neighbor search algorithm collections library written in Rust 🦀 .
★ 2.7ktriton. Development repository for the Triton language and compiler
★ 20ktorch-mlir. The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.
★ 1.9kakg. AKG (Auto Kernel Generator) is an optimizer for operators in Deep Learning Networks, which provides the ability to automatically fuse ops with specific patterns.
★ 258custom_matmul_kernels. Customized matrix multiplication kernels
★ 57OTA. Official implementation of our CVPR2021 paper "OTA: Optimal Transport Assignment for Object Detection" in Pytorch.
★ 246torchshard. Slicing a PyTorch Tensor Into Parallel Shards
★ 300bagua. Bagua Speeds up PyTorch
★ 881bagua-core. Core communication lib for Bagua.
★ 48folly. An open-source C++ library developed and used at Facebook.
★ 30kMaskFormer. Per-Pixel Classification is Not All You Need for Semantic Segmentation (NeurIPS 2021, spotlight)
★ 1.5kmlcc. 一些有用的功能
★ 48tengine-pipe. Tengine 管子是用来快速生产 demo 的辅助工具
★ 11Pretrained-IPT. Python
★ 470cccc-lite. C++
★ 52FisherPruning. Group Fisher Pruning for Practical Network Compression(ICML2021)
★ 164OpenCL-101. Learn OpenCL step by step.
★ 140vision. Datasets, Transforms and Models specific to Computer Vision
★ 2naive-ui. A Vue 3 Component Library. Fairly Complete. Theme Customizable. Uses TypeScript. Fast.
★ 18kForward. A library for high performance deep learning inference on NVIDIA GPUs.
★ 556dino. PyTorch code for Vision Transformers training with the Self-Supervised learning method DINO
★ 7.6k