This is your work, valued
LTX-Video. LTXVideo Q8
★ 115q8_kernels. Cuda
★ 82fastgrpo. Python
★ 26anigraph. anime
★ 2paella_vortrag. DSML KZ Paella
★ 1sicp_sol. SICP SOLUTIONS
★ 1mllib. C++
★ 1konakona666.github.io. CV
★ 1nurbot_detection. ANTINURBOT
★ 1heart. Jupyter Notebook
★ 1ps2env. Python
★ 1game_agent. JavaScript
★ 1cv_prepare_olimp_2026. Jupyter Notebook
★ 1MoonEP. MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
★ 959low-latency-nccl. C++
★ 44gigatoken. Language model tokenization at GB/s
★ 3.8klingbot-vision. Self-supervised learning for spatial perception
★ 888tml-fa4. FA4-based Relative Attention Kernel developed by TML and Colfax
★ 17ShapeR. Code for the ShapeR research paper
★ 864swav. PyTorch implementation of SwAV https//arxiv.org/abs/2006.09882
★ 2.1kkrea-2. Python
★ 22lingbot-video. Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
★ 894SimCLR. PyTorch implementation of SimCLR: A Simple Framework for Contrastive Learning of Visual Representations
★ 2.5kdinov3. Reference PyTorch implementation and models for DINOv3
★ 11kUnifiedReward. Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex
★ 796DanceGRPO. An official implementation of DanceGRPO: Unleashing GRPO on Visual Generation
★ 1.6kHelios. Helios: Real Real-Time Long Video Generation Model
★ 2kgdrcopy. A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology
★ 1.4kMegakernels. Kernels, of the mega variety :)
★ 788Etha. Python
★ 10PipelineRL. A scalable asynchronous reinforcement learning implementation with in-flight weight updates.
★ 430hermes-agent. The agent that grows with you
★ 223kfastgrpo. Python
★ 26lucebox. LLM speculative inference server for consumer hardware & heterogeneous computing
★ 2.7kgemlite. Fast low-bit matmul kernels in Triton
★ 479marlin. FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.
★ 1.1kFlashQLA. high-performance linear attention kernel library built on TileLang
★ 622flashinfer. FlashInfer: Kernel Library for LLM Serving
★ 6.1kOpenClaw-RL. OpenClaw-RL: Train any agent simply by talking
★ 5.6kares. ares is a cross-platform, open source, multi-system emulator, focusing on accuracy and preservation.
★ 1.7ktokenspeed. TokenSpeed is a speed-of-light LLM inference engine.
★ 1.8kSEAL. Self-Adapting Language Models
★ 1.8kcontext-1-data-gen. Python
★ 433Search-R1. Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
★ 5.2kSZ3. Error-bounded Lossy Data Compressor (for floating-point/integer datasets)
★ 123cuSZ. A GPU accelerated error-bounded lossy compression for scientific data.
★ 100flash-attention-residuals. Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)
★ 86cuSZp. Fast GPU error-bounded lossy compressor for floating-point data.
★ 68PLFM_RADAR. Open-source, low-cost 10.5 GHz PLFM phased array RADAR system
★ 23kkraken. Triton-based Symmetric Memory operators and examples
★ 109xccl. C
★ 26DualPipe. A bidirectional pipeline parallelism algorithm for computation-communication overlap in DeepSeek V3/R1 training.
★ 3kDeepEP. DeepEP: an efficient expert-parallel communication library
★ 9.9kfirecracker. Secure and fast microVMs for serverless computing.
★ 36kprime-rl. Agentic RL Training at Scale
★ 1.8klibTAS. GNU/Linux software to (hopefully) give TAS tools to games
★ 583llm-code-editor. LLM Code Editor (think Claude Code)
★ 1KazakhstanOlympiadAI-HomeTask. Python
★ 7transitions. A lightweight, object-oriented finite state machine implementation in Python with many extensions
★ 6.6kml_conformer_generator. A tool for spatially-aware molecule design and optimisation via Equivariant Diffusion and GCN
★ 58flightmare. An Open Flexible Quadrotor Simulator
★ 1.4kColosseum. Open source simulator for autonomous robotics built on Unreal Engine with support for Unity
★ 668freat. A simple game hacking tool, leveraging Godot and Frida for cross-platform support and :star: style points :star:.
★ 6pysc2. StarCraft II Learning Environment
★ 8.3kPhoenixOS. Fast OS-level support for GPU checkpoint and restore
★ 287selkies. Open-Source Low-Latency Accelerated Linux WebRTC HTML5 Remote Desktop Streaming Platform for Self-Hosting, Containers, Kubernetes, or Cloud/HPC
★ 2kqemu. QEMU in a Docker container.
★ 1.9kGymnasium. A standard API for single-agent reinforcement learning environments, with popular reference environments and related utilities (formerly Gym)
★ 12kshardy. MLIR-based partitioning system
★ 201einops. Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)
★ 9.6kscalable-dreamer. Python
★ 10LTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5ktriton-ext. A collection of out-of-tree extensions for the Triton language and compiler
★ 32torchtitan. A PyTorch native platform for training generative AI models
★ 5.6kcomfy-kitchen. Fast kernel library for Diffusion inference with multiple compute backends.
★ 124cutile-python. cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
★ 2.1kDistCA. Efficient Long-context Language Model Training by Core Attention Disaggregation
★ 106sonic-moe. Accelerating MoE with IO and Tile-aware Optimizations
★ 734Duix-Mobile. 🚀 The best real-time interactive AI avatar(digital human) with on-premise deployment and <1.5 s latency.
★ 8.2krocSHMEM. [DEPRECATED] Moved to ROCm/rocm-systems repo
★ 145clice. A next-generation C++ language server for modern C++, focused on high performance and deep code intelligence
★ 1.3knvbench. CUDA Kernel Benchmarking Library
★ 914anny. Anny, A Free and Interpretable Human Body Model for all ages, written in PyTorch.
★ 545HLA. Official Project Page for HLA: Higher-order Linear Attention (https://arxiv.org/abs/2510.27258)
★ 102flash-linear-attention. 🚀 Efficient implementations for emerging model architectures
★ 5.5ktritonparse. TritonParse: A Compiler Tracer, Visualizer, and Reproducer for Triton Kernels
★ 212botorch. Bayesian optimization in PyTorch
★ 3.6kCode2Video. [ICML 2026] Video generation via code
★ 1.9kmulti-gpu-programming-models. Examples demonstrating available options to program multiple GPUs in a single node or a cluster
★ 909llmq. Quantized LLM training in pure CUDA/C++.
★ 252ffpa-attn. FFPA: Kernel Library for Large Headdim Attention - 1.5x~6x speedup over PyTorch SDPA.
★ 319cache-dit. A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
★ 1.2ktriton. Github mirror of trition-lang/triton repo.
★ 182nvshmem. NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
★ 567accelerated-computing-hub. NVIDIA curated collection of educational resources related to general purpose GPU programming.
★ 1.9kflux. A fast communication-overlapping library for tensor/expert parallelism on GPUs.
★ 1.4kmirage. Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
★ 2.4kFPGA-Project-2022-simple-tpu. Systolic array based simple TPU for CNN on PYNQ-Z2
★ 51SimpleTPU. A FPGA Based CNN accelerator, following Google's TPU V1.
★ 176Triton-distributed. Distributed Compiler based on Triton for Parallel Systems
★ 1.5kreal-work. Real Work is a framework for RL environments that simulate real work.
★ 10cccl. CUDA Core Compute Libraries
★ 2.4kLTX-Video. Official repository for LTX-Video
★ 11kring-flash-attention. Ring attention implementation with flash attention
★ 1kpplx-kernels. Perplexity GPU Kernels
★ 595attention-gym. Helpful tools and examples for working with flex-attention
★ 1.2kefficientvit. Efficient vision foundation models for high-resolution generation and perception.
★ 3.3ktuning_playbook. A playbook for systematically maximizing the performance of deep learning models.
★ 30ktorch-dct. DCT (discrete cosine transform) functions for pytorch
★ 641sycl-tla. SYCL* Templates for Linear Algebra (SYCL*TLA) - SYCL based CUTLASS implementation for Intel GPUs
★ 80wmma_tensorcore_sample. Matrix Multiply-Accumulate with CUDA and WMMA( Tensor Core)
★ 148accelerated-model-architectures. Python
★ 91cuda_hgemm. Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruction.
★ 559flow_grpo. [NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL
★ 2.4ktpux. A set of Python scripts that makes your experience on TPU better
★ 56MagiAttention. A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
★ 894ComfyUI_ProfilerX. Node and workflow profiling. Find bottlenecks in your workflows. See trends over time.
★ 82LTX-Video-Trainer. Community trainer for Lightricks' LTX Video model 🎬 ⚡️
★ 466LTX-Video-Q8-Kernels. Python
★ 83FasterTransformer. Transformer related optimization, including BERT, GPT
★ 6.4kComfyUI-LTXVideo. LTX-Video Support for ComfyUI
★ 4kDeepGEMM. DeepGEMM: clean and efficient BLAS kernel library on GPU
★ 7.6kyamajiflow. Python
★ 19PolyDye. Full Color Printer Mod for Marlin 3D Printers
★ 441Hunyuan3D-2. High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
★ 14kSPAN. This repository is an official PyTorch implementation of our paper “Lightweight Image Super-Resolution with Sliding Proxy Attention Network”
★ 7flow_matching. A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
★ 4.7kunsloth. Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.
★ 69kFastSoftmax. Step by step implementation of a fast softmax kernel in CUDA
★ 70terrain-navigation. Implementation for safe low altitude navigation in steep terrain for fixed-wing Aerial Vehicles
★ 186LTX-Video. LTXVideo Q8
★ 115q8_kernels. Cuda
★ 82VideoSys. VideoSys: An easy and efficient system for video generation
★ 2kchannels-last-groupnorm. A CUDA kernel for NHWC GroupNorm for PyTorch
★ 23Parallel-Computing-Cuda-C. CUDA Learning guide
★ 568CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13ktriton. Development repository for the Triton language and compiler
★ 20kstable-fast. https://wavespeed.ai/ Best inference performance optimization framework for HuggingFace Diffusers on NVIDIA GPUs.
★ 1.3kMambaVision. [CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
★ 2.2kNATTEN. Fast Multi-dimensional Sparse Attention
★ 780FALCONN. FAst Lookups of Cosine and Other Nearest Neighbors (based on fast locality-sensitive hashing)
★ 1.2kmra-attention. Python
★ 8cuda-checkpoint. CUDA checkpoint and restore utility
★ 477Open-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kdiffae. Official implementation of Diffusion Autoencoders
★ 969pytorch-binary-converter. Turning float tensors to binary tensors according to IEEE-754 standard.
★ 36jepa. PyTorch code and models for V-JEPA self-supervised learning from video.
★ 4.1kphenaki. A phenaki reproduction using pytorch.
★ 220co-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5kAnimateDiff. Official implementation of AnimateDiff.
★ 12kcommon_metrics_on_video_quality. You can easily calculate FVD, PSNR, SSIM, LPIPS for evaluating the quality of generated or predicted videos.
★ 583Latte. [TMLR 2025] Latte: Latent Diffusion Transformer for Video Generation.
★ 1.9kLatte. The official implementation of Latte: Latent Diffusion Transformer for Video Generation.
★ 34tarproc. Utilities for sequential processing of tar files.
★ 25webdataset. A high-performance Python-based I/O system for large (and small) deep learning problems, with strong support for PyTorch.
★ 3.2kTensor-Puzzles. Solve puzzles. Improve your pytorch.
★ 4.3kagile_autonomy. Repository Containing the Code associated with the Paper: "Learning High-Speed Flight in the Wild"
★ 799QuestPatcher. Generic il2cpp modding tool for Oculus Quest (1/2/3) apps.
★ 379LucidDreamer. Official code for the paper "LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes".
★ 1.5kanimatediff-cli-prompt-travel. animatediff prompt travel
★ 1.2kPASD. [ECCV2024] Pixel-Aware Stable Diffusion for Realistic Image Super-Resolution and Personalized Stylization
★ 1kcutlass. CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10khiggsfield. Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
★ 4kIP-Adapter. The image prompt adapter is designed to enable a pretrained text-to-image diffusion model to generate images with image prompt.
★ 6.6khack-assembler. Compiler for simple assembly language
★ 1LLaMPPL. A domain-specific probabilistic programming language for modeling and inference with language models
★ 143awesome-image-translation. Collection of awesome resources on image-to-image translation.
★ 1.2kHS-Diffusion. HS-Diffusion: Semantic-Mixing Diffusion for Head Swapping
★ 91dcface. Python
★ 155StyleMask. Authors official PyTorch implementation of the "StyleMask: Disentangling the Style Space of StyleGAN2 for Neural Face Reenactment" [FG 2023].
★ 113manifesto. The OpenTF Manifesto expresses concern over HashiCorp's switch of the Terraform license from open-source to the Business Source License (BSL) and calls for the tool's return to a truly open-source license.
★ 36klatent-pose-reenactment. The authors' implementation of the "Neural Head Reenactment with Latent Pose Descriptors" (CVPR 2020) paper.
★ 181e4s. (CVPR 2023) E4S: Fine-grained Face Swapping via Regional GAN Inversion
★ 452alga. Algebraic graphs
★ 759facechain. FaceChain is a deep-learning toolchain for generating your Digital-Twin.
★ 9.5kHyperKohaku. A diffusers based implementation of HyperDreamBooth
★ 139perfusion-pytorch. Implementation of Key-Locked Rank One Editing, from Nvidia AI
★ 238flash-attention. Fast and memory-efficient exact attention
★ 25ksd-scripts. Python
★ 7.2kRecurrentGPT. Official Code for Paper: RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text
★ 998Instruction-Tuning-Papers. Reading list of Instruction-tuning. A trend starts from Natrural-Instruction (ACL 2022), FLAN (ICLR 2022) and T0 (ICLR 2022).
★ 769Awesome-Story-Generation. This repository collects an extensive list of awesome papers about Story Generation / Storytelling, exclusively focusing on the era of Large Language Models (LLMs).
★ 641Awesome-Controllable-T2I-Diffusion-Models. A collection of resources on controllable generation with text-to-image diffusion models.
★ 1.1kawesome-personalization. ✨ A curated list of awesome things related to personalization.
★ 26prompt-in-context-learning. Awesome resources for in-context learning and prompt engineering: Mastery of the LLMs such as ChatGPT, GPT-3, and FlanT5, with up-to-date and cutting-edge updates.
★ 2.2kRETRO-pytorch. Implementation of RETRO, Deepmind's Retrieval based Attention net, in Pytorch
★ 879tree-of-thought-llm. [NeurIPS 2023] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
★ 6kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24ktrl. Train transformer language models with reinforcement learning.
★ 19kViCo. Official PyTorch codes for the paper: "ViCo: Detail-Preserving Visual Condition for Personalized Text-to-Image Generation"
★ 242llm-numbers. Numbers every LLM developer should know
★ 4.3kfastcomposer. [IJCV] FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
★ 715Stable-Diffusion-Inpaint. Stable diffusion for inpainting
★ 229Awesome-Transformer-Attention. An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites
★ 5.1ktomesd. Speed up Stable Diffusion with this one simple trick!
★ 1.4kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kLLM-As-Chatbot. LLM as a Chatbot Service
★ 3.3kPAIR-Diffusion. [CVPR 2024] PAIR Diffusion: A Comprehensive Multimodal Object-Level Image Editor
★ 522multimodal-garment-designer. This is the official repository for the paper "Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing". ICCV 2023
★ 445TurkicASR. A multilingual ASR model that can recognize ten Turkic languages—Azerbaijani, Bashkir, Chuvash, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Uyghur, and Uzbek.
★ 87FixNoise. Official Pytorch Implementation for "Fix the Noise: Disentangling Source Feature for Controllable Domain Translation" (CVPR 2023, CVPRW 2022 Best paper)
★ 174DreamArtist-stable-diffusion. stable diffusion webui with contrastive prompt tuning
★ 866Text2Poster-ICASSP-22. Official implementation of the ICASSP-2022 paper "Text2Poster: Laying Out Stylized Texts on Retrieved Images"
★ 214sd_personalization_encoder. Unofficial implementation of Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models
★ 24ELITE. ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation (ICCV 2023, Oral)
★ 541TaskMatrix. Python
★ 34kGLIGEN. Open-Set Grounded Text-to-Image Generation
★ 2.2kSSHarmonization. [ICCV'2021] "SSH: A Self-Supervised Framework for Image Harmonization", Yifan Jiang, He Zhang, Jianming Zhang, Yilin Wang, Zhe Lin, Kalyan Sunkavalli, Simon Chen, Sohrab Amirghodsi, Sarah Kong, Zhangyang Wang
★ 102DeepImageBlending. This is a Pytorch implementation of deep image blending
★ 471sparsednn. Fast sparse deep learning on CPUs
★ 55ProDiff. PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
★ 432Empathy-Mental-Health. Repository containing codes and dataset access instructions for the EMNLP 2020 paper on empathy in text-based mental health support
★ 188Paella. Official Implementation of Paella https://arxiv.org/abs/2211.07292v2
★ 748