This is your work, valued
Master student at ICT, studying computer vision and generative models.
ctrlora. [ICLR 2025] Codebase for "CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation"
★ 268hitsz-minirv-1. HITSZ 2020春 计算机设计与实践课程,实现基于 miniRV-1 的单周期和流水线CPU
★ 42diffusion-models-pytorch. Implement Diffusion Models with PyTorch.
★ 28HCL. Codebase of ICCV 2023 paper "Hierarchical Contrastive Learning for Pattern-Generalizable Image Corruption Detection"
★ 25mathematical-modeling-python. Python codes for mathematical modeling.
★ 13maskgit-pytorch. Unofficial PyTorch implementation of MaskGIT.
★ 10hitsz-maillab. HITSZ 2022春 计算机网络实验课程,实现一个简单邮件客户端
★ 8gans-pytorch. Implement GANs with PyTorch.
★ 7flow-matching-pytorch. Implement Flow Matching with PyTorch.
★ 7hitsz-net-lab-2022. HITSZ 2022春 计算机网络实验课程,实现一个简单协议栈
★ 6xv6-mit-6.S081-2020.
★ 6CUMCM-2021-A. CUMCM 2021 A,广东省一等奖
★ 5competitive-programming-template. Collection of code templates for Competitive Programming.
★ 5visual-tokenizer-pytorch. Implement visual tokenizers with PyTorch.
★ 4image-backbones-pytorch. Implement image backbones with PyTorch.
★ 4discdiff. Improving Diffusion Models with Discriminative Objectives.
★ 4easy-control. Quickly try out various controllable text-to-image diffusion models with a few lines of code!
★ 3vaes-pytorch. Implement VAEs with PyTorch.
★ 2datasetx. Implements commonly used datasets based on torch and torchvision.
★ 2ign-pytorch. Unofficial implementation of Idempotent Generative Network.
★ 2dlutils. Deep Learning Utilities (PyTorch).
★ 1hitsz-fuse-filesystem. HITSZ 2021秋 操作系统实验五,实现基于 FUSE 的青春版 EX2 文件系统
★ 1hitsz-dip-2022. HITSZ 2022春 数字图像处理 projects
★ 1comment-analysis. 大一立项——结合用户评价分析的网店假货预警
★ 1drifting-models-pytorch. Unofficial PyTorch implementation of "Generative Modeling via Drifting".
★ 1openvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7kocto. Octo is a transformer-based robot policy trained on a diverse mix of 800k robot trajectories.
★ 1.7kSelf-Flow. [ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
★ 676awesome-vla-wam. A Curated List of Vision-Language-Action (VLA) and World Action Models (WAM) Research and Beyond
★ 878mira. Code for MIRA: Multiplayer Interactive World Models with Representation Autoencoders
★ 467minit2i-torch. Official PyTorch re-implementation of MiniT2I.
★ 290UniVideo. [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos
★ 543Echo-Memory. A Simple Baseline for Video World Models with Memory
★ 239vggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14krcm. rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
★ 774WEM. Python
★ 18mochi. The best OSS video generation models, created by Genmo
★ 3.7kHunyuanVideo-1.5. HunyuanVideo-1.5: A leading lightweight video generation model
★ 4.5kcodex. Lightweight coding agent that runs in your terminal
★ 102kPiD. PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
★ 1kkairos. Official code for world model Kairos
★ 2.4kHiDream-O1-Image. Python
★ 1.5kSkyReels-V3. SkyReels V3: Multimodal Video Generation Model
★ 524SkyReels-V2. SkyReels-V2: Infinite-length Film Generative model
★ 7.3kLongLive. Long Video Gen Infrastructure
★ 2.5kRAEv2. Official Implemenation for RAEv2: Improved Baselines with Representation Autoencoders
★ 310minWM. A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
★ 750Awesome-Video-World-Models-with-AR-Diffusion. A Curated List of Awesome Video World Models with AR Diffusion: Covering Algorithms, Applications, and Infrastructure, Aimed at Serving as a Comprehensive Resource for Researchers, Practitioners, and Enthusiasts.
★ 678CausVid. (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
★ 1.4kCausal-Forcing. [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
★ 888FD-Loss. Python
★ 549SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4ktuna-2. Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
★ 739decord. An efficient video loader for deep learning with smart shuffling that's super easy to digest
★ 2.5kpatch-forcing. [CVPR 2026] Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
★ 101Astra. [ICLR 2026] Astra : General Interactive World Model with Autoregressive Denoising"
★ 340YUME. The official code of Yume
★ 679lingbot-world. Advancing Open-source World Models
★ 4.3ktrash-cli. Command line interface to the freedesktop.org trashcan.
★ 4.4kPixelDiT. [CVPR 2026 Best Paper Finalist] Pixel Diffusion Transformers for Image Generation
★ 904GRN. [ECCV 2026] Generative Refinement Networks for Visual Synthesis (Support C2I & T2I & T2V)
★ 145flow-maps. Official codebase for the paper "How to build a consistency model: Learning flow maps via self-distillation" (NeurIPS 2025).
★ 135Adversarial-Flow-Models. Python
★ 85py-meanflow. Pytorch implementation for MeanFlow
★ 363MeanFlow. Pytorch implementation of MeanFlow on ImageNet and CIFAR10
★ 504stylegan-t. [ICML'23] StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis
★ 1.2kJoyAI-Image. JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
★ 2.2kbtop. A monitor of resources
★ 34kedge_eval_python. A python implementation of edge eval
★ 69UNITE-tokenization-generation. Single-stage End-to-End Training for Tokenization and Generation
★ 117gmmn. Generative moment matching networks
★ 152drifting. Python
★ 483CubiD. [CVPR2026 Highlight] Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens https://arxiv.org/abs/2603.19232
★ 63coulomb_gan. Implementation of Coulomb GANs
★ 64drifting-model. Personal PyTorch implementation of "Generative Modeling via Drifting" with Claude
★ 240MMD-GAN. MMD-GAN: Towards Deeper Understanding of Moment Matching Network
★ 203BitDance. BitDance & UniWeTok: Open-source autoregressive model with binary visual tokens. A research project for building powerful multimodal autoregressive model.
★ 480flash-bidirectional-linear-attention. Triton implement of bi-directional (non-causal) linear attention
★ 78DiG. [CVPR 2025] DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
★ 185PixelGen. Official repository for “PixelGen: Improving Pixel Diffusion with Perceptual Loss”
★ 275BAR. [ICML 2026] code & model for arxiv paper "Autoregressive Image Generation with Masked Bit Modeling"
★ 60imeanflow. Official Implementation of iMF https://arxiv.org/abs/2512.02012
★ 333pMF. Official Implementation of pMF https://arxiv.org/abs/2601.22158
★ 270FastGen. NVIDIA FastGen: Fast Generation from Diffusion Models
★ 882iFSQ. iFSQ & LlamaGen-REPA
★ 103Scale-RAE. Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
★ 255Self-Transcendence. [ECCV 2026] Official code repository for "Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?"
★ 37SpeedrunDiT. SR-DiT Speedrunning ImageNet Diffusion
★ 139MoGe. [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
★ 2.7kRollingForcing. [ICLR 2026] Official Repo for Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
★ 449GameGen-X.
★ 341MaskDiT. Code for Fast Training of Diffusion Models with Masked Transformers
★ 429Meissonic. [ICLR 2025] Official Implementation of Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
★ 345Muon. Muon is an optimizer for hidden layers in neural networks
★ 2.7kpixel-perfect-depth. [NeurIPS 2025] Pixel-Perfect Depth
★ 1.1kGLM-Image. GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
★ 1kSiD. PyTorch code and model checkpoints for Score identity Distillation (SiD) and its adversarial version (SiDA)
★ 156REPA-E. [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers
★ 512nepa. PyTorch implementation of NEPA
★ 339FramePack. Lets make video diffusion practical!
★ 17kNextFlow. NextFlow🚀: Unified Sequential Modeling Activates Multimodal Understanding and Generation
★ 331Internal-Guidance. CVPR 2026 (Highlight)-Guiding a Diffusion Transformer with the Internal Dynamics of Itself (IG)
★ 84Grounded-SAM-2. Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2
★ 3.7kpy360convert. Python implementation of convertion between equirectangular, cubemap and perspective. (equirect2cube, cube2equirect, equirect2perspec)
★ 601ctm. Python
★ 326GFT. Python
★ 53SoFlow. [ICLR 2026] SoFlow: Solution Flow Models for One-Step Generative Modeling
★ 161CoP. The Collapse of Patches
★ 58htop. htop - an interactive process viewer
★ 8.2kfd. A simple, fast and user-friendly alternative to 'find'
★ 44kHY-WorldPlay. HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency
★ 1.6kFlowEdit. Official implementation of the paper: "FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models"
★ 1kGRAG-Image-Editing. https://little-misfit.github.io/GRAG-Image-Editing/
★ 119torchdiffeq. Differentiable ODE solvers with full GPU support and O(1)-memory backpropagation.
★ 6.5kSVG-T2I. [Arxiv 2025] Official PyTorch Implementation of "SVG-T2I: Scaling up Text-to-Image Latent Diffusion Model Without Variational Autoencoder".
★ 152JoPano. JoPano: Unified Panorama Generation via Joint Modeling
★ 24TwinFlow. [ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻
★ 537SFD. [CVPR 2026] Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
★ 153Awesome-Pixel-Flow.
★ 38REG. [NeurIPS 2025 Oral] Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
★ 274progressive_growing_of_gans. Progressive Growing of GANs for Improved Quality, Stability, and Variation
★ 6.2kZ-Image. Python
★ 12kDeCo. [CVPR2026 Highlight] Official repository for “DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation”
★ 238flux2. Official inference repo for FLUX.2 models
★ 2.6kkandinsky-5. Kandinsky 5.0: A family of diffusion models for Video & Image generation
★ 803clash-verge-rev. A modern GUI client based on Tauri, designed to run in Windows, macOS and Linux for tailored proxy experience
★ 134kimproved-aesthetic-predictor. CLIP+MLP Aesthetic Score Predictor
★ 1.3kAwesome-Visual-Tokenizer. Awesome Visual Tokenizers/Autoencoders
★ 20JiT. PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
★ 2.5kDDT. [CVPR 2026] DDT: Decoupled Diffusion Transformer
★ 408EPG. [ICLR2026] There is No VAE: End-To-End Pixel-Space Generative Modeling Via Self-Supervised Pre-Training
★ 152DanceGRPO. An official implementation of DanceGRPO: Unleashing GRPO on Visual Generation
★ 1.6kEditScore. [ICLR 2026] EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
★ 256Edit-R1. Edit-R1: Reinforce Image Editing with Diffusion Negative-Aware Finetuning and MLLM Implicit Feedback
★ 295pico-banana-400k. Python
★ 1.8kGPT-Image-Edit. GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
★ 243UniControl. Unified Controllable Visual Generation Model
★ 662zsh-z. Jump quickly to directories that you have visited "frecently." A native Zsh port of z.sh with added features.
★ 2.4kSVG. [ICLR 2026] Official PyTorch Implementation of "Latent Diffusion Model Without Variational Autoencoder".
★ 457HunyuanImage-3.0. HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
★ 3.2kFSQ-pytorch. A Pytorch Implementation of Finite Scalar Quantization
★ 191ImageReward. [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7kHPSv2. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
★ 677RAE. Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
★ 2kReDi. [NeurIPS'25 Spotlight] Boosting Generative Image Modeling via Joint Image-Feature Synthesis
★ 121ULMEvalKit. ULMEvalKit: One-Stop Eval ToolKit for Image Generation
★ 56termcolor. ANSI color formatting for output in terminal
★ 325clean-fid. PyTorch - FID calculation with proper image resizing and quantization steps [CVPR 2022]
★ 1.2kIQA-PyTorch. 🔎 🖼️ 🔥PyTorch Toolbox for Image Quality Assessment, including PSNR, SSIM, LPIPS, FID, NIQE, NRQM(Ma), MUSIQ, TOPIQ, NIMA, DBCNN, BRISQUE, PI and more...
★ 3.3krich. Rich is a Python library for rich text and beautiful formatting in the terminal.
★ 57kUni-X. [ICLR 2026] Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
★ 10Lumina-DiMOO. Lumina-DiMOO - An Open-Sourced Multi-Modal Large Diffusion Language Model
★ 1kST-AR.
★ 14geneval. GenEval: An object-focused framework for evaluating text-to-image alignment
★ 472DiffusionNFT. [ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
★ 994DCGAN-LSGAN-WGAN-GP-DRAGAN-Tensorflow-2. Reimplementation of GANs
★ 424SRPO. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
★ 1.3kbpytop. Linux/OSX/FreeBSD resource monitor
★ 11klogos. A huge collection of SVG logos
★ 6.8kTiM. Transition Models
★ 156NExT-GPT. Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6kInternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kVLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kSEED-X. Multimodal Models in Real World
★ 558HiDream-I1. Python
★ 2.5kHiDream-E1. Python
★ 790MagicQuill. [CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System
★ 3.7kNextStep-1. [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s Multimodal Intelligence team.
★ 693FreeNoise. [ICLR 2024] Code for FreeNoise based on VideoCrafter
★ 428fx. Terminal JSON viewer & processor
★ 21kSRA. [ICLR 2026] Self-Representation Alignment for Diffusion Transformers (SRA)
★ 147DispLoss. Python
★ 110REPA. [ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
★ 1.7kplatonic-rep. Python
★ 715Wan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kdinov3. Reference PyTorch implementation and models for DINOv3
★ 11kUniPic. Open-source SOTA multi-image editing model
★ 871HPSv3. Official implementation of HPSv3: Towards Wide-Spectrum Human Preference Score (ICCV2025)
★ 331Tar. [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
★ 202Qwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kPixNerd. [ICLR 2026] PixNerd: Pixel Neural Field Diffusion
★ 185ml-inrflow. Python
★ 73HunyuanWorld-1.0. Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
★ 2.9kSDEdit. PyTorch implementation for SDEdit: Image Synthesis and Editing with Stochastic Differential Equations
★ 1.2kGigaTok. [ICCV 2025] Official repo for "GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation"
★ 204radial-attention. [NeurIPS 2025] Radial Attention: O(nlogn) Sparse Attention with Energy Decay for Long Video Generation
★ 605supervision. We write your reusable computer vision tools. 💜
★ 48kFUDOKI. [NeurIPS 2025 Spotlight] FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
★ 77ReVQ. Explore how to get a VQ-VAE models efficiently!
★ 69Kimi-K2. Kimi K2 is the large language model series developed by Moonshot AI team
★ 11kaac-datasets. Audio Captioning datasets for PyTorch.
★ 128Ovis-U1. An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image editing within a single powerful framework.
★ 450Step1X-Edit. A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.
★ 2.2kx-flux. Python
★ 2.2kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kAwesome-Unified-Understanding-and-Generation.
★ 52AnyGPT. Code for "AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling"
★ 880Awesome-Unified-Multimodal-Models. Awesome Unified Multimodal Models
★ 1.3kmetaquery. Official Implementation of Paper Transfer between Modalities with MetaQueries
★ 325FluxKits. Python
★ 110meanflow. JAX implementation of MeanFlow
★ 602UniFork. UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
★ 48OmniFlows. The official implementation of OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
★ 137OmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kMixture-of-Transformers. Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR 2025.
★ 281Ming. Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.
★ 665DeltaFM. [ICCV 2025] Official Implementation of Contrastive Flow Matching
★ 184SimpleAR. Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"
★ 431UniWorld. UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
★ 886Chain-of-Zoom. [NeurIPS'25 Spotlight] Official repository for "Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment"
★ 777eqvae. [ICML'25] EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling.
★ 182CFG-Zero-star. Official repo for CFG-Zero*
★ 715Diffusion-LLM-Papers. A Collection of Papers on Diffusion Language Models
★ 181LLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kMuddit. [ICLR 2026] Official Implementation of Muddit [Meissonic II]: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model.
★ 119DDO. [ICML 2025 Spotlight] Direct Discriminative Optimization: Reinforcing Diffusion/Autoregressive with GAN Discrimination
★ 124Harmon. [ICCV2025]Code Release of Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
★ 192Jodi. Jodi: Unification of Visual Generation and Understanding via Joint Modeling
★ 92BLIP3o. Official implementation of BLIP3o-Series
★ 1.7kLLaDA-V. Python
★ 348pytorch-normalizing-flows. Normalizing flows in PyTorch. Current intended use is education not production.
★ 917MMaDA. MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
★ 1.7k