This is your work, valued
BeyondTimeShifts. Beyond Time Shifts (ECCV 2026): a reference-free audio-visual synchronization evaluator built on Qwen2.5-Omni-3B, with SynthSync data, R-GRPO training, and the SyncBench benchmark.
★ 6JavisDiT. [ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.
★ 377r1_reward. ✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
★ 292OmniReward. Python
★ 48LLaVA-Hound-DPO. Python
★ 159Omni-RRM. Python
★ 1perception_models. State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
★ 2.3kTalkVid. TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis [CVPR 2026 Findings]
★ 194DreamOmni2. This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
★ 2kinsightface. State-of-the-art 2D and 3D Face Analysis Project
★ 29kJoyAI-Echo. JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation
★ 1.8kDreamID-Omni. [ICML 2026] DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
★ 275Wan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kAudio-Omni. [SIGGRAPH 2026] Repository of Audio-Omni
★ 400ID-LoRA. [ECCV 2026] Generate high resolution videos with a custom voice and appearance, based on LTX-2/LTX-2.3 + Identity In-Context LoRA
★ 349OmniCustom. Official Implementation of 'OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model'
★ 426Phantom. Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
★ 1.5kIdentity-as-Presence. Python
★ 16Qwen3-Omni. Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
★ 3.9kConsisID. [CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
★ 848ROLL. An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3.3kal-folio. A beautiful, simple, clean, and responsive Jekyll theme for academics
★ 16kmellea. Mellea is a library for writing generative programs.
★ 1.8kAwesome-Video-Generation-Post-Training. [TMLR] Video Generation Models: A Survey of Post-Training and Alignment | 🔥 A continuously updated collection of papers, datasets, and benchmarks on post-training and alignment for video generation.
★ 175python3-cookbook. 《Python Cookbook》 3rd Edition Translation
★ 12kDiffusionNFT. [ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
★ 999TIIF-Bench. Official repository for the paper "TIIF-Bench: How Does Your T2I Model Follow Your Instructions?".
★ 128OneIG-Benchmark. [NeurIPS 2025 DB] OneIG-Bench is a meticulously designed comprehensive benchmark framework for fine-grained evaluation of T2I models across multiple dimensions, including subject-element alignment, text rendering precision, reasoning-generated content, stylization, and diversity.
★ 121VACE. [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
★ 3.9kSpatialT2I. [CVPR 2026🔥] Enhancing Spatial Understanding in Image Generation via Reward Modeling
★ 86Pref-GRPO. Official implementation of Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
★ 276Flow-Factory. A unified framework for easy reinforcement learning in Flow-Matching models
★ 645EditReward. EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing [ICLR 2026]
★ 156UnifiedReward. Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex
★ 796InternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kVLM-R1. Solve Visual Understanding with Reinforced VLMs
★ 6kUniPercept. [ICML2026 Spotlight] UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
★ 159LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74klightly-studio. Curate, Annotate, and Manage Your Data in LightlyStudio.
★ 871Awesome-Unified-Multimodal-Models. Awesome Unified Multimodal Models
★ 1.3kCLIPood. About Code Release for "CLIPood: Generalizing CLIP to Out-of-Distributions" (ICML 2023), https://arxiv.org/abs/2302.00864
★ 70SRPO. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
★ 1.3kDanceGRPO. An official implementation of DanceGRPO: Unleashing GRPO on Visual Generation
★ 1.6kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kQwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kGLM-V. GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 2.4klearn-user-pref. Official implementation of "Learning User Preferences for Image Generation Models"
★ 14MoMu. The PyTorch implementation of MoMu, described in "Natural Language-informed Modeling of Molecule Graphs".
★ 29Enformer_Borzoi_Training_Pytorch. We provide pytorch training scripts for enformer and borzoi
★ 10aesthetic-predictor-v2-5. SigLIP-based Aesthetic Score Predictor
★ 427SPACE. Python
★ 20MindtheGap. ICCV25 highlight
★ 59acad-homepage.github.io. AcadHomepage: A Modern and Responsive Academic Personal Homepage
★ 2.9kDAS. Official implementation for Diffusion Alignment as Sampling (DAS), ICLR'25, Spotlight
★ 66Bagel. Open-source unified multimodal model
★ 6.1kflow_grpo. [NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL
★ 2.4kddpo. Code for the paper "Training Diffusion Models with Reinforcement Learning"
★ 574DeepSeek-Math. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
★ 3.4kDual-Diffusion. Code for D-DiT
★ 69ParetoFlow. [ICLR 2025] Official Implementation of ParetoFlow: Guided Flows in Multi-Objective Optimization🧬🧬🧬
★ 29Awesome-Multi-Objective-Deep-Learning. A comprehensive list of gradient-based multi-objective optimization algorithms in deep learning.
★ 114Janus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18klatent-consistency-model. Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
★ 4.6krg-lcd. Reward Guided Latent Consistency Distillation
★ 26MetaMask. Python
★ 13CPKP. Python
★ 4Uniform-Attention-Maps. [WACV 2025] Uniform Attention Maps: Enhancing Image Fidelity in Reconstruction and Editing
★ 17Reward-Instruct. [NeurIPS 2025] Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation
★ 35CrossFlow. [CVPR2025] PyTorch-based reimplementation of CrossFlow, as proposed in 'Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution'
★ 345transfusion-pytorch. Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
★ 1.4kmar. PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838
★ 1.9kTACO. TACO: TFBS-Aware Cis-Regulatory Element Optimization
★ 23ProtDETR. Interpretable Enzyme Function Prediction via Residue-Level Detection
★ 13newPCMDM. Python
★ 13PCMDM. Official Implementation of "Synthesizing Long-Term Human Motions with Diffusion Models via Coherent Sampling"
★ 14GraphTextRetrieval. Part of official implementation of "Natural language-informed learning of molecule graphs"
★ 18InstantStyle. InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation 🔥
★ 2kZegCLIP. Official implement of CVPR2023 ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation
★ 258Fk-Diffusion-Steering. A general framework for inference-time scaling and steering of diffusion models with arbitrary rewards.
★ 230SimpleTuner. A general fine-tuning kit geared toward image/video/audio diffusion models.
★ 2.9kgeneval. GenEval: An object-focused framework for evaluating text-to-image alignment
★ 473AlignProp. AlignProp uses direct reward backpropogation for the alignment of large-scale text-to-image diffusion models. Our method is 25x more sample and compute efficient than reinforcement learning methods (PPO) for finetuning Stable Diffusion
★ 324PickScore. Python
★ 601BLIP. PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
★ 5.7kImageReward. [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7kopen-prompts. Python
★ 790textcraftor. Python
★ 14TexForce. Official PyTorch codes for "Enhancing Diffusion Models with Text-Encoder Reinforcement Learning", ECCV2024
★ 59ELLA. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
★ 1.3ktext-to-text-transfer-transformer. Code for the paper "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer"
★ 6.5kT2I-CompBench. [Neurips 2023 & TPAMI] T2I-CompBench (++) for Compositional Text-to-image Generation Evaluation
★ 345sd3.5. Python
★ 1.5kReNO. [NeurIPS 2024] ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
★ 166Structured-Diffusion-Guidance. Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
★ 321CogVLM2. GPT4V-level open-source multi-modal model based on Llama3-8B
★ 2.4kOpen-O1. Python
★ 1.3kopen-o1. open-o1: Using GPT-4o with CoT to Create o1-like Reasoning Chains
★ 116mm-cot. Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)
★ 4kaesthetic-predictor. A linear estimator on top of clip to predict the aesthetic quality of pictures
★ 728NExT-GPT. Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6kRPG-DiffusionMaster. [ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
★ 1.8kPMG. The repository of paper Personalized Multimodal Response Generation with Large Language Models
★ 18Vitron. NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
★ 576long_stable_diffusion. Long-form text-to-images generation, using a pipeline of deep generative models (GPT-3 and Stable Diffusion)
★ 692SEED. Official implementation of SEED-LLaMA (ICLR 2024).
★ 642gill. 🐟 Code and models for the NeurIPS 2023 paper "Generating Images with Multimodal Language Models".
★ 470HPSv2. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
★ 677ComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kflux. Application Architecture for Building User Interfaces
★ 17kDEADiff. [CVPR 2024] Official implementation of "DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations"
★ 280google-research. Google Research
★ 38kdaam. Diffusion attentive attribution maps for interpreting Stable Diffusion.
★ 803stable-diffusion-webui-daam. DAAM for Stable Diffusion Web UI
★ 177LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kAttend-and-Excite. Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models" (SIGGRAPH 2023)
★ 770training-free-structured-diffusion-guidance. 🤗 Unofficial huggingface/diffusers-based implementation of the paper "Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis".
★ 120SeedSelect. Code for our papers : "Generating images of rare concepts using pre-trained diffusion models" (AAAI 24) and "Norm-guided latent space exploration for text-to-image generation" (Neurips 23)
★ 87StyleDrop-PyTorch. Unoffical implement for [StyleDrop](https://arxiv.org/abs/2306.00983)
★ 588sd-webui-fabric. Python
★ 408fabric. Python
★ 319initno. [CVPR 2024] InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization
★ 80direct-preference-optimization. Reference implementation for DPO (Direct Preference Optimization)
★ 2.9kOpen-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kCogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kMCTS-DPO. This is the repository that contains the source code for the Self-Evaluation Guided MCTS for online DPO.
★ 331ziplora-pytorch. Implementation of "ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs"
★ 565SPO. [CVPR 2025] Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
★ 271mapo. Official codebase for Margin-aware Preference Optimization for Aligning Diffusion Models without Reference (MaPO).
★ 83VADER. Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as VideoCrafter, OpenSora, ModelScope and StableVideoDiffusion by finetuning them using various reward models such as HPS, PickScore, VideoMAE, VJEPA, YOLO, Aesthetics etc.
★ 316d3po. [CVPR 2024] Code for the paper "Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model"
★ 244DiffusionDPO. Code for "Diffusion Model Alignment Using Direct Preference Optimization"
★ 707Free-Guidance-Diffusion. A novel method that provides greater control over generated images by guiding the internal representations of the pre-trained Stable Diffusion.
★ 40PAE. [CVPR 2024] Dynamic Prompt Optimizing for Text-to-Image Generation
★ 87LLaMA2-Accessory. An Open-source Toolkit for LLM Development
★ 2.8kIP-Adapter. The image prompt adapter is designed to enable a pretrained text-to-image diffusion model to generate images with image prompt.
★ 6.6kDreambooth-Stable-Diffusion. Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) by way of Textual Inversion (https://arxiv.org/abs/2208.01618) for Stable Diffusion (https://arxiv.org/abs/2112.10752). Tweaks focused on training faces, objects, and styles.
★ 3.2kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kDreambooth-Stable-Diffusion. Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
★ 7.7kCLIP-Adapter. Python
★ 580Transfer-Learning-Library. Transfer Learning Library for Domain Adaptation, Task Adaptation, and Domain Generalization
★ 3.9kBatchFormer. CVPR2022, BatchFormer: Learning to Explore Sample Relationships for Robust Representation Learning, https://arxiv.org/abs/2203.01522
★ 253APE. [ICCV 2023] Code for "Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior Refinement"
★ 150TPT. Test-time Prompt Tuning (TPT) for zero-shot generalization in vision-language models (NeurIPS 2022))
★ 214CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kICLR24. Official code for ICLR 2024 paper, "A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation"
★ 86SuS-X. Code for the paper: "SuS-X: Training-Free Name-Only Transfer of Vision-Language Models" [ICCV'23]
★ 104Tip-Adapter. Python
★ 677DPT. Official PyTorch implementation for "Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels"
★ 96VAEs. Variational autoencoders: VAE, gaussian mixture VAE (GMVAE), and a basic ladder VAE (LVAE)
★ 47AugSelf. Python
★ 46self-supervised. Whitening for Self-Supervised Representation Learning | Official repository
★ 137DomainBed. DomainBed is a suite to test domain generalization algorithms
★ 1.6klightly. A python library for self-supervised learning on images.
★ 3.8kLINKX_paddle. Unofficial PaddlePaddle Implementation of LINKX (Large Scale Learning on Non-Homophilous Graphs)
★ 3