This is your work, valued
PhDing...
Awesome-Open-Vocabulary. (TPAMI 2024) A Survey on Open Vocabulary Learning
★ 999DiffSensei. Implementation of [CVPR 2025] "DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation"
★ 923MotionBooth. [NeurIPS 2024 Spotlight] The official implement of research paper "MotionBooth: Motion-Aware Customized Text-to-Video Generation"
★ 138Language-Driven-Video-Inpainting. (CVPR 2024) Official code for paper "Towards Language-Driven Video Inpainting via Multimodal Large Language Models"
★ 99betrayed-by-captions. (ICCV 2023) Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation
★ 48robust-ref-seg. (TIP 2024) Towards Robust Referring Image Segmentation
★ 40Does-Hearing-Help-Seeing. Python
★ 19download-cc3m. Jupyter Notebook
★ 5Disentangled-Clustering-Contrastive-Learning-of-Time-Series-Representation. Jupyter Notebook
★ 2cosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kUniCharacter. An official implementation of Towards Customized Multimodal Role-Play.
★ 10LLaDA2.0-Uni. LLaDA2.0-Uni: Understanding and Generation the World.
★ 770Kiwi-Edit. A unified and fully open-source framework for instruction-guided and reference-guided video editing using natural language.
★ 311MOVA. MOVA: Towards Scalable and Synchronized Video–Audio Generation
★ 1.1kLTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5kSora2-mini. UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
★ 57HunyuanVideo-1.5. HunyuanVideo-1.5: A leading lightweight video generation model
★ 4.5kmolmo2. Code for the Molmo2 Vision-Language Model
★ 696UniVideo. [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos
★ 543Does-Hearing-Help-Seeing. Python
★ 19GPT-Image-Edit. GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
★ 243BLIP3o. Official implementation of BLIP3o-Series
★ 1.7kHunyuanImage-3.0. HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
★ 3.2kemma. Official repo for paper "EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture."
★ 62DraCo. Offical Repository for Paper: DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
★ 18tuna.
★ 94AIA. Python
★ 45Step1X-Edit. A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.
★ 2.2kMirage. [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
★ 294Monet. [CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"
★ 216Emu3.5. Native Multimodal Models are World Learners
★ 1.5kUniLIP. [ICLR 2026 🔥 ] Official implementation of "UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing"
★ 151Qwen3-Omni. Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
★ 3.9kAnyV2V. Code and data for "AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks" [TMLR 2024]
★ 655VACE. [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
★ 3.9kQwen2-Audio. The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 2.1kQwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
★ 4.1kUniVerse-1-code. The official UniVerse-1 code.
★ 129Koala-36M. Official implementation of the paper "Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content".
★ 253Waver. Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.
★ 950UniCTokens. A framework for unified personalized model, achieving mutual enhancement between personalized understanding and generation. Demonstrating the potential of cross-task information transfer in personalized scenario, paving the way for the development of general unified models.
★ 131UniPic. Open-source SOTA multi-image editing model
★ 871memfof. [ICCV'2025 Highlight] MEMFOF: High-Resolution Training for Memory-Efficient Multi-Frame Optical Flow Estimation
★ 102thinking-with-generated-images. Doodling our way to AGI ✏️ 🖼️ 🧠
★ 128tarsier. Tarsier -- a family of large-scale video-language models, which is designed to generate high-quality video descriptions , together with good capability of general video understanding.
★ 548X-Omni. Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).
★ 427ARC-Hunyuan-Video-7B. Structured Video Comprehension of Real-World Shorts
★ 240VideoLLaMA2. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
★ 1.3kKeye. Python
★ 809Wan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kCtrl-Crash. Official PyTorch Implementation of Ctrl-Crash 💥
★ 53AudioLDM2. Text-to-Audio/Music Generation
★ 2.6kIndex-anisora. Python
★ 2.5kstar-vector. StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
★ 4.5kCoDi. CoDi:Subject-Consistent and Pose-Diverse Text-to-Image Generation
★ 36Lumos. [ICLR 2026] Lumos Project: Frontier video unified model research by Alibaba DAMO Academy.
★ 161Ovis-U1. An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image editing within a single powerful framework.
★ 450attention-interpolation-diffusion. [NeurIPS 2024] Official Implementation of Attention Interpolation of Text-to-Image Diffusion
★ 110ControlNet. Generate videos that interpolate between two given images
★ 101ShareGPT-4o-Image. Python
★ 285VMoBA. Official implementation of paper "VMoBA: Mixture-of-Block Attention for Video Diffusion Models"
★ 64OmniCorpus. [ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
★ 425TPDM. Implementation of "Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation" [CVPR 2025]
★ 43JavisDiT. [ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.
★ 376MMaDA. MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
★ 1.7kBagel. Open-source unified multimodal model
★ 6.1kPSDiffusion.
★ 18TeaCache. Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
★ 1.4kFastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kSparse-VideoGen. [ICML2025, NeurIPS2025 Spotlight] Sparse VideoGen 1 & 2: Accelerating Video Diffusion Transformers with Sparse Attention
★ 697DiTFastAttn. Jupyter Notebook
★ 192MAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.7kSpargeAttn. [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
★ 1krlhf-flow-diffusion. Reinforcement Learning from Human Feedback for Flow-Based Diffusion Models
★ 6ddpo-pytorch. DDPO for finetuning diffusion models, implemented in PyTorch with LoRA support
★ 768Wan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kHunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kCogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kMoBA. MoBA: Mixture of Block Attention for Long-Context LLMs
★ 2.2kAwesome-Visual-Autoregressive-Model. Latest Advances on Autoregressive Visual Models.📖
★ 28hart. HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
★ 647NATTEN. Fast Multi-dimensional Sparse Attention
★ 779DeepSeek-V3. Python
★ 104kGameFactory. [ICCV 2025] GameFactory: Creating New Games with Generative Interactive Videos
★ 495CLEAR. [NeurIPS 2025] Official PyTorch implementation of paper "CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up".
★ 219LinFusion. Official PyTorch and Diffusers Implementation of "LinFusion: 1 GPU, 1 Minute, 16K Image"
★ 317VSSD. [ICCV2025] Introduce Mamba2 to Vision.
★ 190tread. Python
★ 182Vim. [ICML 2024] Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
★ 3.9kSa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.7kLTX-Video. Official repository for LTX-Video
★ 11kmamba. Mamba SSM architecture
★ 19kAniDoc. [CVPR'25] Official Implementations for Paper - AniDoc: Animation Creation Made Easier
★ 573U-DiT. [NeurIPS 2024] The official code of "U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers"
★ 240SakugaDataset. Official Repository for Sakuga-42M Dataset
★ 81anomaly-seg. The Combined Anomalous Object Segmentation (CAOS) Benchmark
★ 158dataset-api. The ApolloScape Open Dataset for Autonomous Driving and its Application.
★ 619diamond. DIAMOND (DIffusion As a Model Of eNvironment Dreams) is a reinforcement learning agent trained in a diffusion world model. NeurIPS 2024 Spotlight.
★ 2.1kVBench. [CVPR2024 Highlight] VBench - We Evaluate Video Generation
★ 1.7kSana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kGameGen-X.
★ 341Playable-Game-Generation. An open-source lightweight game generation paradigm. It includes everything from data processing to model architecture design and playability-based evaluation methods. The game runs at 20 FPS on a single consumer-grade graphics card (RTX-2060) while maintaining high playability.
★ 120ditflow. Official PyTorch implementation - Video Motion Transfer with Diffusion Transformers
★ 81DiffSensei. Implementation of [CVPR 2025] "DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation"
★ 923physgen. PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation (ECCV 2024)
★ 352Emu. Emu Series: Generative Multimodal Models from BAAI
★ 1.8kSEED-X. Multimodal Models in Real World
★ 558Show-o. [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
★ 2kCSGO. CSGO: Content-Style Composition in Text-to-Image Generation 🔥
★ 390StyleID. [CVPR 2024 Highlight] Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models for Style Transfer
★ 481X-Pose. [ECCV 2024] Official implementation of the paper "X-Pose: Detecting Any Keypoints"
★ 815DeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kInstantID. Python
★ 154InstantID. InstantID: Zero-shot Identity-Preserving Generation in Seconds 🔥
★ 12kinsightface. State-of-the-art 2D and 3D Face Analysis Project
★ 29kawesome-comics-understanding. The official repo of the Comics Survey: "A missing piece in Vision and Language: A Survey on Comics Understanding"
★ 139Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kQwen-VL. The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
★ 6.7kInternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kUniPortrait. [ICCV2025] UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization
★ 275VideoTetris. [NeurIPS 2024] VideoTetris: Towards Compositional Text-To-Video Generation
★ 236direct_a_video. Python
★ 95gradio. Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!
★ 43kSEED-Story. SEED-Story: Multimodal Long Story Generation with Large Language Model
★ 884ReVersion. [SIGGRAPH Asia 2024] ReVersion: Diffusion-Based Relation Inversion from Images
★ 503Analogist. Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model (SIGGRAPH 2024)
★ 38MS-Diffusion. [ICLR 2025] Official implementation of MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
★ 311ai-comic-factory. Generate comic panels using a LLM + SDXL. Powered by Hugging Face 🤗
★ 1.3kIP-Adapter. The image prompt adapter is designed to enable a pretrained text-to-image diffusion model to generate images with image prompt.
★ 6.6kStoryGen. [CVPR 2024] Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models
★ 270TheaterGen. TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation
★ 69AutoStudio. [CVPRW 2026] AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
★ 452ToonCrafter. [SIGGRAPH Asia 2024, Journal Track] ToonCrafter: Generative Cartoon Interpolation
★ 6kMG-LLaVA. Official repository for paper MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning(https://arxiv.org/abs/2406.17770).
★ 160MotionBooth. [NeurIPS 2024 Spotlight] The official implement of research paper "MotionBooth: Motion-Aware Customized Text-to-Video Generation"
★ 138mangadex. A python wrapper for the mangadex API V5. Work in progress
★ 46magi. Generate a transcript for your favourite Manga: Detect manga characters, text blocks and panels. Order panels. Cluster characters. Match texts to their speakers. Perform OCR.
★ 460ComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kmanga109api. Simple python API to read annotation data of Manga109
★ 130Omost. Your image is almost there!
★ 7.6kstyle-aligned. Official code for "Style Aligned Image Generation via Shared Attention"
★ 1.3kcross-image-attention. Officail Implementation for "Cross-Image Attention for Zero-Shot Appearance Transfer"
★ 404InST. Official implementation of the paper “Inversion-Based Style Transfer with Diffusion Models” (CVPR 2023)
★ 588FreeStyle. FreeStyle : Free Lunch for Text-guided Style Transfer using Diffusion Models
★ 132CameraCtrl. Python
★ 658MotionEditor. [CVPR2024] MotionEditor is the first diffusion-based model capable of video motion editing.
★ 188Moore-AnimateAnyone. Character Animation (AnimateAnyone, Face Reenactment)
★ 3.5kchamp. [ECCV 2024] Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
★ 4.3kMagicDance. [ICML 2024] MagicPose(also known as MagicDance): Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion
★ 777DreamPose. Official implementation of "DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion"
★ 1kDisCo. [CVPR2024] DisCo: Referring Human Dance Generation in Real World
★ 1.1kFIFO-Diffusion_public. Official implementation of FIFO-Diffusion: Generating Infinite Videos from Text without Training (NeurIPS 2024)
★ 486handy_voting. handy tools for user study
★ 21AnimateDiff. Official implementation of AnimateDiff.
★ 12kFreeCustom. [CVPR 2024] Official PyTorch implementation of FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
★ 177StoryDiffusion. Accepted as [NeurIPS 2024] Spotlight Presentation Paper
★ 6.4kPanda-70M. [CVPR 2024] Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
★ 701VideoBooth. [CVPR2024] VideoBooth: Diffusion-based Video Generation with Image Prompts
★ 309VGen. Official repo for VGen: a holistic video generation ecosystem for video generation building on diffusion models
★ 3.2kFreeNoise. [ICLR 2024] Code for FreeNoise based on VideoCrafter
★ 428VAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kSEINE. [ICLR 2024] SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
★ 967StreamingT2V. [CVPR 2025] StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
★ 1.6kDenseDiffusion. Official Pytorch Implementation of DenseDiffusion (ICCV 2023)
★ 508StableCascade. Official Code for Stable Cascade
★ 6.5kReferFormer. [CVPR2022] Official Implementation of ReferFormer
★ 356ovsam. [ECCV 2024] The official code of paper "Open-Vocabulary SAM".
★ 1kOMG-Seg. Official Repo For OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]
★ 1.4kTrailBlazer. [SIGGRAPH Asia 2024] TrailBlazer: Trajectory Control for Diffusion-Based Video Generation
★ 102MotionCtrl. Official Code for MotionCtrl [SIGGRAPH 2024]
★ 1.5kPhotoMaker. PhotoMaker [CVPR 2024]
★ 10kxformers. Hackable and optimized Transformers building blocks, supporting a composable construction.
★ 11kPeekaboo. Interactive Video Generation via Masked-Diffusion
★ 110VideoCrafter. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
★ 5.1kmodelscope. ModelScope: bring the notion of Model-as-a-Service to life.
★ 9.1kLaVie. [IJCV 2024] LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models
★ 952Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kGen-L-Video. The official implementation for "Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising".
★ 308BoxDiff. [ICCV 2023] BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion
★ 275freecontrol. Official implementation of CVPR 2024 paper: "FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition"
★ 480tianxingwu.github.io. My homepage https://tianxingwu.github.io
★ 5FreeInit. [ECCV 2024] FreeInit: Bridging Initialization Gap in Video Diffusion Models
★ 544ElasticDiffusion-official. The official Pytorch Implementation for ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation (CVPR 2024)
★ 160MultiDiffusion. Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
★ 1.1kVideoDirectorGPT. official implementation of VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning (COLM 2024)
★ 181Paint-by-Example. Paint by Example: Exemplar-based Image Editing with Diffusion Models
★ 1.3kLLM-groundedDiffusion. LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models (LLM-grounded Diffusion: LMD, TMLR 2024)
★ 483LLM-groundedVideoDiffusion. [ICLR 2024] LLM-grounded Video Diffusion Models (LVD): official implementation for the LVD paper
★ 172Mix-of-Show. NeurIPS 2023, Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models
★ 428AutoStory. [IJCV'24] AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort
★ 149UDiffText. [ECCV 2024] Official repo for UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models
★ 236generative-models. Generative Models by Stability AI
★ 27kacad-homepage.github.io. AcadHomepage: A Modern and Responsive Academic Personal Homepage
★ 2.9kshields. Concise, consistent, and legible badges in SVG and raster format
★ 27kPixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kReLA. [CVPR 2023 Highlight & IJCV 2026] GRES: Generalized Referring Expression Segmentation
★ 689ContextDET. Contextual Object Detection with Multimodal Large Language Models
★ 261groundingLMM. [CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.
★ 964lang-segment-anything. SAM with text prompt
★ 2.6kinstruct-pix2pix. Python
★ 6.9kdst-det. [TCSVT] state-of-the-art open vocabulary detector on COCO/LVIS/V3Det
★ 35Mini-DALLE3. Mini-DALLE3: Interactive Text to Image by Prompting Large Language Models
★ 313LISA. Project Page for "LISA: Reasoning Segmentation via Large Language Model"
★ 2.7kODISE. Official PyTorch implementation of ODISE: Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models [CVPR 2023 Highlight]
★ 945ProPainter. [ICCV 2023] ProPainter: Improving Propagation and Transformer for Video Inpainting
★ 6.8kindex-X. 禁书目录X系列,不一样的阅读体验!
★ 648InstructDiffusion. PyTorch implementation of InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions.
★ 445Inpaint-Anything. Inpaint anything using Segment Anything and inpainting models.
★ 7.7k