This is your work, valued
MobileOne. An Improved One millisecond Mobile Backbone
★ 147Few-Shot-Adversarial-Learning-for-face-swap. This is a unofficial re-implementation of the paper "Few-Shot Adversarial Learning of Realistic Neural Talking Head Models"
★ 139Face_Recognition_using_pytorch. Using MTCNN and MobileFaceNet on Face Recognition
★ 67Morph-UGATIT. a morph transfer UGATIT for image translation.
★ 61FaceID-GAN. this is a re-implementation of CVPR2018 paper "FaceID-GAN"
★ 39HandAI. Using hand gestures control different effect of photography.
★ 35FlowNet1.0-using-Keras. This model's weights are converted from Flownet of Nvidia
★ 12pytorch3d. PyTorch3D is FAIR's library of reusable components for deep learning with 3D data
★ 1initial-GAN-tutorial-On-Mnist. Generative Adversarial Nets base on Tensorflow and mnist
★ 1Bernini. Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
★ 1.2klingbot-world-v2. Infinite Worlds with Versatile Interactions
★ 1.4kXiaomi-Robotics-1. Code for Xiaomi-Robotics-1
★ 273RTDMD. [arXiv 2026] This is the official PyTorch implementation of "RTDMD: Reinforcing Few-step Generators via Reward-Tilted Distribution Matching".
★ 41PFM. Official implementation of "Perceptual Flow Matching for Few-Step Generative Modeling"
★ 11DanceOPD. 🔥 DanceOPD: On-Policy Generative Field Distillation
★ 341convert_to_quant. Python
★ 146patch-forcing. [CVPR 2026] Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
★ 99DreamX-World. DreamX-World: A General-Purpose Interactive World Model
★ 735gamecraft-bench. Code and Data for paper "GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?"
★ 187minit2i-torch. Official PyTorch re-implementation of MiniT2I.
★ 290Qwen-Image-Edit-Causal. In our implementation of Qwen-Image-Edit, we employ block causal attention to improve inference speed.
★ 54RiT. PyTorch implementation of RiT: Vanilla Diffusion Transformers Suffice in Representation Space
★ 27LakonLab. Official implementation of AsymFlow, pi-Flow, GMFlow
★ 457PiD. PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
★ 997UniCharacter. An official implementation of Towards Customized Multimodal Role-Play.
★ 10WritingAIPaper. Writing AI Conference Papers: A Handbook for Beginners
★ 3.9k1.x_distill. The official code repo of 1.x-Distill, is a stagewise distillation framework for diversity, high-quality and efficient few-step generation, with support for fractional-step inference and MLP-based cache acceleration.
★ 18daily_stock_analysis. LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
★ 59ktuna-2. Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
★ 739PAE. Official Implementation of "What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion"
★ 80sapiens2. 1K resolution vision transformers pretrained on 1B human images.
★ 884torchsde. Differentiable SDE solvers with GPU support and efficient sensitivity analysis.
★ 1.7kFlow-Factory. A unified framework for easy reinforcement learning in Flow-Matching models
★ 643Lance. A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
★ 1.3kadvantage_weighted_matching. Official code for paper Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
★ 94SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4kHiDream-O1-Image. Python
★ 1.5kD-OPSD. Official Repo of "D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models"
★ 291TwinFlow. [ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻
★ 537SlowFast. PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
★ 7.4kdiffusion-pipe. A pipeline parallel training script for diffusion models.
★ 2kai-toolkit. The ultimate training toolkit for finetuning diffusion models
★ 11kSimpleTuner. A general fine-tuning kit geared toward image/video/audio diffusion models.
★ 2.9kUniPercept. [ICML2026 Spotlight] UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
★ 158ArtiMuse. [CVPR 2026] ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding(书生 · 妙析多模态美学理解大模型)
★ 215World-R1. [ICML 2026] World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
★ 411Nucleus-Image. NucleusImage training recipe
★ 83InfraTech. 分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
★ 3.2kFastGen. NVIDIA FastGen: Fast Generation from Diffusion Models
★ 873ODTSR. [CVPR2026] ODTSR: This repo is the official implementation of "One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution"
★ 184MegaStyle. MegaStyle, 面向一致性与多样性的可扩展风格数据生成框架
★ 131sphere-encoder. PyTorch Implementation of Image Generation with a Sphere Encoder
★ 44LLaDA2.0-Uni. LLaDA2.0-Uni: Understanding and Generation the World.
★ 770OpenMythos. A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
★ 15kFireRed-Image-Edit. FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity consistency, and seamless multi-element fusion.
★ 1.3kAnthropomorphicIntelligence. Advancing AI by embracing human-likeness for better AI understanding, human–AI collaboration, and social simulation, bridging technology and genuine human experience.
★ 132ai4animationpy. A Python framework for AI-driven character animation using neural networks.
★ 2.1kchinese-poetry. The most comprehensive database of Chinese poetry 🧶最全中华古诗词数据库, 唐宋两朝近一万四千古诗人, 接近5.5万首唐诗加26万宋诗. 两宋时期1564位词人,21050首词。
★ 53ksscd-copy-detection. Open source implementation of "A Self-Supervised Descriptor for Image Copy Detection" (SSCD).
★ 419AdaRefSR. AdaRefSR is a novel reference-based one-step diffusion super-resolution framework. Paper was accepted by ICLR2026.
★ 67Awesome-Nano-Banana-images. A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the community's development of image generation and unified models(click to website to see our blog)
★ 23kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 92kHelios. Helios: Real Real-Time Long Video Generation Model
★ 2kT2ITrainer. Practice Code for text to image trainer
★ 559pillow_heif. Python library for working with HEIF images and plugin for Pillow.
★ 302face-parsing. Real-time face parsing and facial semantic segmentation with BiSeNet - PyTorch training, ONNX export, pretrained weights.
★ 305SegFace. [AAAI 25] SegFace: Face Segmentation of Long-tail classes
★ 105PixelGen. Official repository for “PixelGen: Improving Pixel Diffusion with Perceptual Loss”
★ 275rt_gene. RT-GENE: Real-Time Eye Gaze and Blink Estimation in Natural Environments
★ 443dino_perceptual. DINO-based perceptual losses and FDD feature extraction
★ 37WeDLM. WeDLM: The fastest diffusion language model with standard causal attention and native KV cache compatibility, delivering real speedups over vLLM-optimized baselines.
★ 648jekyll-theme-chirpy. A minimal, responsive, and feature-rich Jekyll theme for technical writing.
★ 10kLTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5kface_anon_simple. [WACV 2025] Official implementation of "Face Anonymization Made Simple"
★ 212TurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kcosmos-predict2.5. Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the form of video.
★ 1.3krcm. rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
★ 774nepa. PyTorch implementation of NEPA
★ 339GenEval2. Evaluation codes and data for GenEval2
★ 80sam-audio. The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 3.6kdacvae. DACVAE
★ 226VTP. [ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generation
★ 496aesthetic-predictor-v2-5. SigLIP-based Aesthetic Score Predictor
★ 426HPSv3. Official implementation of HPSv3: Towards Wide-Spectrum Human Preference Score (ICCV2025)
★ 331Z-Image. Python
★ 12kPhantom. Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
★ 1.5kHuMo. HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
★ 1.3kidentity-grpo. Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
★ 203FaceCLIP. Python
★ 42DSPO. Python
★ 46awesome-alignment-of-diffusion-models. [ACM Computing Surveys] The collection of awesome papers on alignment of diffusion models.
★ 430Bagel. Open-source unified multimodal model
★ 6.1kDCM. [ICCV2025] DCM: Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
★ 206Echo-4o. Echo-4o: Harnessing Proprietary Models’ Synthetic Images for Improved Image Generation
★ 506qwen-image-finetune. Repo for Qwen Image Finetune
★ 249FSDP-Training. Minimal PyTorch implementation of TP, SP, FSDP and sharded-EMA
★ 32FastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kDiffSynth-Studio. Enjoy the magic of Diffusion models!
★ 13kyolo-face. YOLO Face 🚀 in PyTorch
★ 877DC-Gen. DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
★ 400SDAR. SDAR (Synergy of Diffusion and AutoRegression), a large diffusion language model(1.7B, 4B, 8B, 30B)
★ 364flymyai-lora-trainer. Qwen-Image text to image lora trainer
★ 764pixel3dmm. [Official Code] Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
★ 512Ditto. [CVPR'26 Highlight] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
★ 618OmniVinci. OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 675DDGS. Official implementation of DDGS: Depth-and-Density Guided Gaussian Splatting for Stable and Accurate Sparse-View Reconstruction
★ 86RAE. Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
★ 2knanochat. The best ChatGPT that $100 can buy.
★ 57kUAE. Official repository for the UAE paper, unified-GRPO, and unified-Bench
★ 166Show-o. [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
★ 2kToonComposer. [ICLR 2026] Streamlining Cartoon Production with Generative Post-Keyframing
★ 585dinov3. Reference PyTorch implementation and models for DINOv3
★ 11kPusa-VidGen. Pusa: Thousands Timesteps Video Diffusion Model
★ 686TransformerEngine. A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
★ 3.5kUnifiedReward. Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex
★ 796DanceGRPO. An official implementation of DanceGRPO: Unleashing GRPO on Visual Generation
★ 1.6kOpenRLHF. An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
★ 9.9kDeTok. Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"
★ 195ddpo-pytorch. DDPO for finetuning diffusion models, implemented in PyTorch with LoRA support
★ 768CharaConsist. Official implementation of ICCV 2025 paper - CharaConsist: Fine-Grained Consistent Character Generation
★ 166OmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kSana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6knunchaku. [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
★ 3.9kSeedVR2. [ICLR2026] SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
★ 825Phantom-Data. Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset
★ 117Jenga. [NeurIPS 2025] Training-Free Efficient Video Generation via Dynamic Token Carving
★ 287flow_grpo. [NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL
★ 2.4kICEdit. [NeurIPS 2025] Image editing is worth a single LoRA! 0.1% training data for fantastic image editing! Surpasses GPT-4o in ID persistence~ MoE ckpt released! Only 4GB VRAM is enough to run!
★ 2.1kMing. Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.
★ 665SkyReels-A2. SkyReels-A2: Compose anything in video diffusion transformers
★ 713perception_models. State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
★ 2.3kUniRig. [SIGGRAPH 2025] One Model to Rig Them All: Diverse Skeleton Rigging with UniRig
★ 1.7kSimpleAR. Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"
★ 431MAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.7kREPA-E. [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers
★ 511LLMDet. (CVPR 2025 highlight✨) Official repository of paper "LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models"
★ 608SAM2MOT. [AAAI 2026] Code for "SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation".
★ 180sapiens. High-resolution models for human tasks.
★ 5.4kFramePack. Lets make video diffusion practical!
★ 17kOmniPaint. [ICCV 25] OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting
★ 328UNO. [ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
★ 1.4kSegAnyMo. [CVPR 2025] Code for Segment Any Motion in Videos
★ 485LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kRF-Solver-Edit. [🚀ICML 2025] "Taming Rectified Flow for Inversion and Editing" Using FLUX and HunyuanVideo for image and video editing!
★ 638LLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kInstruct-CLIP. Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning (CVPR 2025)
★ 34Awesome-Inference-Time-Scaling. Paper List of Inference/Test Time Scaling/Computing
★ 399MagicID.
★ 35star-vector. StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
★ 4.5kMUG-VOS. Official Implementation of "Multi-Granularity Video Object Segmentation" (AAAI 2025)
★ 25SEED-Voken. SEED-Voken: A Series of Powerful Visual Tokenizers
★ 1kNAR. [ICCV 2025] The official implementation of "Neighboring Autoregressive Modeling for Efficient Visual Generation"
★ 62vggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kLeffa. [CVPR 2025] Learning Flow Fields in Attention for Controllable Person Image Generation
★ 1.7kMS-Diffusion. [ICLR 2025] Official implementation of MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
★ 311In-Context-LoRA. Official repository of In-Context LoRA for Diffusion Transformers
★ 2.1kEasyControl. Implementation of "EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer"(ICCV2025)
★ 1.7ktheLMbook. This is the official repository for The Hundred-Page Language Models Book by Andriy Burkov
★ 2.2kStep-Video-T2V. Python
★ 3.2kdeep-q-trading-agent. Deep RL stock trading agent
★ 57VideoAlign. [NeurIPS 2025] Improving Video Generation with Human Feedback
★ 489TinyZero. Minimal reproduction of DeepSeek R1-Zero
★ 13kMotionClone. [ICLR 2025] Official implementation of MotionClone: Training-Free Motion Cloning for Controllable Video Generation
★ 516MatchAnything. Code for "MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training", Arxiv 2025.
★ 1.3kSa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.7ksofttopk. differentiable top-k operator
★ 23direct-preference-optimization. Reference implementation for DPO (Direct Preference Optimization)
★ 2.9kflow-matching. Flow Matching implemented in PyTorch
★ 121X-Pose. [ECCV 2024] Official implementation of the paper "X-Pose: Detecting Any Keypoints"
★ 815DiffSensei. Implementation of [CVPR 2025] "DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation"
★ 923flow_matching. A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
★ 4.6kTrack-Anything. Track-Anything is a flexible and interactive tool for video object tracking and segmentation, based on Segment Anything, XMem, and E2FGVI.
★ 7kP3M-Net. The official repo for [IJCV'23] "Rethinking Portrait Matting with Privacy Preserving"
★ 113OminiControl. [ICCV 2025 Highlight] OminiControl: Minimal and Universal Control for Diffusion Transformer
★ 1.9kDino_V2. Dino V2 for Classification, PCA Visualization, Instance Retrival: https://arxiv.org/abs/2304.07193
★ 206IOPaint. Image inpainting tool powered by SOTA AI Model. Remove any unwanted object, defect, people from your pictures or erase and replace(powered by stable diffusion) any thing on your pictures.
★ 23kaddit. Python
★ 390Trident. [ICCV2025] Harnessing CLIP, DINO and SAM for Open Vocabulary Segmentation
★ 126MagicQuill. [CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System
★ 3.7kmask-grounding. [CVPR2024] Mask Grounding for Referring Image Segmentation
★ 29PixWizard. [ICLR2025] A versatile image-to-image visual assistant, designed for image generation, manipulation, and translation based on free-from user instructions.
★ 211OmniGen. OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
★ 4.3kJanus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kLightRAG. [EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation"
★ 38kEmu3. Next-Token Prediction is All You Need
★ 2.4kCATdiffusion. Official implementation of ImprovingText-guided ObjectInpainting with SemanticPre-inpainting in ECCV 2024
★ 62flux. Official inference repo for FLUX.1 models
★ 26kinvertAvatar. [SIGGRAPH 2024] InvertAvatar: Incremental GAN Inversion for Generalized Head Avatars
★ 606DRepNet. Official Pytorch implementation of 6DRepNet: 6D Rotation representation for unconstrained head pose estimation.
★ 666UnSAM. [NeurIPS 2024] Code release for "Segment Anything without Supervision"
★ 503ml-4m. 4M: Massively Multimodal Masked Modeling
★ 1.8kMG-LLaVA. Official repository for paper MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning(https://arxiv.org/abs/2406.17770).
★ 160SLD. 🔥 [CVPR2024] Official implementation of "Self-correcting LLM-controlled Diffusion Models (SLD)
★ 187chameleon. Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
★ 2.1kres-adapter. [AAAI 2025] Official codes of "ResAdapter: Domain Consistent Resolution Adapter for Diffusion Models".
★ 760AlphaCLIP. [CVPR 2024] Alpha-CLIP: A CLIP Model Focusing on Wherever You Want
★ 876IC-Light. More relighting!
★ 8.5kToonCrafter. [SIGGRAPH Asia 2024, Journal Track] ToonCrafter: Generative Cartoon Interpolation
★ 6kAwesome-Diffusion-Model-Based-Image-Editing-Methods. Diffusion Model-Based Image Editing: A Survey (TPAMI 2025)
★ 712ChatTTS. ChatTTS is a generative speech model for daily dialogue.
★ 1DMD2. (NeurIPS 2024 Oral 🔥) Improved Distribution Matching Distillation for Fast Image Synthesis
★ 1.4kRPG-DiffusionMaster. [ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
★ 1.8ktimesfm. TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
★ 27kMatcher. [ICLR'24 & IJCV‘25] Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
★ 572coyo-dataset. COYO-700M: Large-scale Image-Text Pair Dataset
★ 1.3kArc2Face. [ECCV 2024 Oral 🔥] Arc2Face: A Foundation Model for ID-Consistent Human Faces ------------------------ [ICCVW 2025] ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
★ 803StyleID. [CVPR 2024 Highlight] Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models for Style Transfer
★ 481break-a-scene. Official implementation for "Break-A-Scene: Extracting Multiple Concepts from a Single Image" [SIGGRAPH Asia 2023]
★ 525Face-diffuser. [CVPR2024] Official implementation of High-fidelity Person-centric Subject-to-Image Synthesis.
★ 53