This is your work, valued
Image-Super-Resolution-via-Iterative-Refinement. Unofficial implementation of Image Super-Resolution via Iterative Refinement by Pytorch
★ 3.9kPalette-Image-to-Image-Diffusion-Models. Unofficial implementation of Palette: Image-to-Image Diffusion Models by Pytorch
★ 1.8kdistributed-pytorch-template. This is a seed project for distributed PyTorch training, which was built to customize your network quickly
★ 160Image-Zooming-Using-Directional-Cubic-Convolution-Interpolation. Unofficial Python implementation about Image Zooming Using Directional Cubic Convolution Interpolation (DCC) by Numpy.
★ 18Cminus-Compiler. 【Yacc】Cminus Compiler
★ 10A-Demo-for-Image-Inpainting-by-React. This is a demo for image inpainting project by Flask and React
★ 8face-mask-segmentation. 【Tools】Face recognition and attribute segmentation using Python, dlib, and One Millisecond Face Alignment with an Ensemble of Regression Trees](http://www.cv-foundation.org/openaccess/content_cvpr_2014/papers/Kazemi_One_Millisecond_Face_2014_CVPR_paper.pdf).
★ 5Collaborative-Drawboard. 【SpringBoot】A collaboration drawing board using Fabric and SocketIO communication, and build the database platform
★ 4janspiry.github.io. SCSS
★ 4rewriting. Rewriting a Deep Generative Model, ECCV 2020 (oral). Interactive tool to directly edit the rules of a GAN to synthesize scenes with objects added, removed, or altered. Change StyleGANv2 to make extravagant eyebrows, or horses wearing hats.
★ 3MobileClass. 【Java】Mobile Class System
★ 3MNIST-PyTorch. 【Pytorch】LeNet5 implement for image classification task on MNIST dataset by PyTorch
★ 2Janspiry. Processing
★ 2multivariate_normal. Scipy.stats.multivariate_normal.pdf() implementation by C++ with OpenCV.
★ 2mmrotate. OpenMMLab Rotated Object Detection Toolbox and Benchmark
★ 1ml2021_group. Python
★ 1accio-cu. Hybrid computer use for AI agents on macOS.
★ 11RDM. Python
★ 79S1-Omni-Image. A Unified Multimodal Model for Scientific Image Understanding and Generation
★ 9rich. Rich is a Python library for rich text and beautiful formatting in the terminal.
★ 57kDeepPaperNote. DeepPaperNote is an agent skill for deep-reading a single paper and generating high-quality Obsidian-style research notes. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
★ 566dailypaper-skills. 用Agent skills打造我的论文流水线
★ 1.1kDanceOPD. 🔥 DanceOPD: On-Policy Generative Field Distillation
★ 360DeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kPiD. PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
★ 1kSeedVR2. [ICLR2026] SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
★ 826HYPIR. Official implementation of HYPIR: Harnessing Diffusion-Yielded Score Priors for Image Restoration (SIGGRAPH 2025)
★ 1.2kMARBLE. Multi-Aspect Reward Balance for Diffusion RL
★ 39Boogu-Image. Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.
★ 868i1. Code release for "i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models"
★ 255UniRL. UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
★ 864cosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kD-OPSD. Official Repo of "D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models"
★ 292FaceID-6M. Python
★ 54OPD. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
★ 863OPSD. Python
★ 522Flow-OPD. Official Repo of "Flow-OPD: On-Policy Distillation for Flow Matching Models"
★ 267VAREdit. Python
★ 106HiDream-O1-Image. Python
★ 1.5kPRISM. Beyond SFT-to-RL: Pre-alignment via Black-BoxOn-Policy Distillation for Multimodal RL
★ 98SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4kSenseNova-SI. [CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Models
★ 293Flow-Factory. A unified framework for easy reinforcement learning in Flow-Matching models
★ 644CausVid. (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
★ 1.4kwool_scripts. 收集一些Loon、Surge、QuantumultX、ShadowRocket、Egern的配置与去广告规则。
★ 5.4kProxyResource. 可莉的Loon资源库 | 插件 | 脚本 | 规则
★ 4.8kScaleEdit-12M. ScaleEdit-12M is the largest open-source image editing dataset to date, spanning 23 task families across diverse real and synthetic domains.
★ 14PixelSmile. PixelSmile: Fine-grained facial expression editing with continuous control, reduced semantic entanglement, and strong identity preservation.
★ 475FIRM-Reward. Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation
★ 40SpatialEdit. [Official Repo] SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
★ 215Logics-Parsing. Python
★ 1.4kJoyAI-Image. JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
★ 2.2kOneReward. Python
★ 348GEditBench_v2. GEditBench v2: A Human-Aligned Benchmark for General Image Editing
★ 61isc2021. Code for the Image similarity challenge.
★ 205DeQA-Score. [CVPR 2025] Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution
★ 244Q-Align. ③[ICML2024] [IQA, IAA, VQA] All-in-one Foundation Model for visual scoring. Can efficiently fine-tune to downstream datasets.
★ 613Self-Flow. [ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
★ 684Q-Insight. Q-Insight Family: Q-Insight, VQ-Insight and RALI (NeurIPS 2025 Spotlight, AAAI 2026 Oral, and ICLR 2026 Oral)
★ 314Sa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.7kEditHF. EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
★ 23VeloEdit. Velocity Field Analysis and Intervention for Image Edit
★ 14Stand-In. [CVPR2026 🎉] Stand-In is a lightweight, plug-and-play framework for identity-preserving video generation.
★ 779cleanrl. High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
★ 10kSageAttention. [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
★ 3.5ktianshou. An elegant PyTorch deep reinforcement learning library.
★ 11kTextPecker. [CVPR2026] TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
★ 56pua. 你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.
★ 19kWildActor. Accepted by ICML2026
★ 92SpatialT2I. [CVPR 2026🔥] Enhancing Spatial Understanding in Image Generation via Reward Modeling
★ 86NoHumansRequired. NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
★ 9awesome-openclaw-usecases. A community collection of OpenClaw use cases for making life easier.
★ 32kFireRed-Image-Edit. FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity consistency, and seamless multi-element fusion.
★ 1.3kSViMo_code. This is the repository that contains source code for the [SViMo](https://Droliven.github.io/SViMo_project).
★ 18Reward-Forcing. [CVPR 2026 Highlight] Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
★ 352lingbot-world. Advancing Open-source World Models
★ 4.3kPaper2Video. Automatic Video Generation from Scientific Papers
★ 2.3kMuseTalk. MuseTalk: Real-Time High Quality Lip Synchorization with Latent Space Inpainting
★ 6.3kMusePose. MusePose: a Pose-Driven Image-to-Video Framework for Virtual Human Generation
★ 2.7kArtiMuse. [CVPR 2026] ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding(书生 · 妙析多模态美学理解大模型)
★ 215Think-Then-Generate. HTML
★ 116GLM-Image. GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
★ 1kLTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5kNextFlow. NextFlow🚀: Unified Sequential Modeling Activates Multimodal Understanding and Generation
★ 331DreamOmni3. This project is the official implementation of 'DreamOmni3: Scribble-based Editing and Generation''
★ 40MagicQuillV2. Official Implementations for Paper - MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
★ 153PICABench. PICABench: How Far Are We from Physically Realistic Image Editing?
★ 39UltraHR-100k. This is the official repository of UltraHR-100K.
★ 45T2I-Distill. [Tutorial] Few-Step Distillation for Text-to-Image Generation: A Practical Guide
★ 370Subjects200K. Subjects200K dataset
★ 132ImageCritic. Official implementation of ImageCritic (CVPR 2026)
★ 168motion-edit. Official Repository of paper: "MotionEdit: Benchmarking and Learning Motion-Centric Image Editing"
★ 67RealGen. [ECCV 2026 Oral] RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards.
★ 337EditThinker. Unlocking Iterative Reasoning for Any Image Editor
★ 112MICo-150K. Official repository for the paper "MICo-150K: A Comprehensive Dataset for Multi-Image Composition".
★ 111WiseEdit. Python
★ 16UnicBench. [CVPR 2026] UnicEdit-10M and UnicBench project
★ 42OpenSubject. Python
★ 55TwinFlow. [ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻
★ 537donut. Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
★ 6.9kLongCat-Image. Python
★ 714LeX-Art. Official Implementation of "LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis"
★ 85Z-Image. Python
★ 12kdmdr. [ECCV 2026] Official Code of "Distribution Matching Distillation Meets Reinforcement Learning"
★ 288LLaMA2-Accessory. An Open-source Toolkit for LLM Development
★ 2.8kflux2. Official inference repo for FLUX.2 models
★ 2.6kray. Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
★ 43kLightX2V. Lightweight Image Video Action Generation Inference Framework
★ 2.5kPosterCraft. [ICLR'26] Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework
★ 541LLaDA-V. Python
★ 348LLaDA-1.5.
★ 55LLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kLakonLab. Official implementation of AsymFlow, pi-Flow, GMFlow
★ 457UnifiedReward. Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex
★ 796FG-CLIP. New generation of CLIP with strong fine grained discrimination capability, ICML2026 and ICML2025
★ 765SRPO. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
★ 1.3kAwesome-Flow-RL-Papers. A collection of paper/projects that trains flow matching model/policies via RL.
★ 408json_repair. Repair malformed JSON from LLMs, APIs, logs, and user input in Python.
★ 5.1kEmu3.5. Native Multimodal Models are World Learners
★ 1.5kvariational-diffusion-models. PyTorch implementation of Variational Diffusion Models.
★ 106Edit-R1. Edit-R1: Reinforce Image Editing with Diffusion Negative-Aware Finetuning and MLLM Implicit Feedback
★ 295DiffusionNFT. [ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
★ 998identity-grpo. Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
★ 203EditScore. [ICLR 2026] EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
★ 256Self-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kopen-images-dataset. Open Images is a dataset of ~9 million images that have been annotated with image-level labels and bounding boxes spanning thousands of classes.
★ 1.1kpico-banana-400k. Python
★ 1.8kSelf-Forcing-Plus. Unofficial extension implementation of Self-Forcing to support I2V && 14B training.
★ 381HunyuanWorld-1.0. Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
★ 2.9kDCM. [ICCV2025] DCM: Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
★ 206FastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kSenseFlow. 🚀 [ICLR 2026] SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation
★ 114DMD2. (NeurIPS 2024 Oral 🔥) Improved Distribution Matching Distillation for Fast Image Synthesis
★ 1.4kWithAnyone. ✨ [ICLR'26] WithAnyone is capable of generating high-quality, controllable, and ID consistent images
★ 572xlora. X-LoRA: Mixture of LoRA Experts
★ 280Omni-Effects. [AAAI2026] Implementation Code for Omni-Effects
★ 175MOELoRA-peft. [SIGIR'24] The official implementation code of MOELoRA.
★ 193sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kMinerU. Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
★ 76kPaddleOCR. Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
★ 87kstyle-tokenizer. Python
★ 112OmniStyle. OmniStyle: Filtering High Quality Style Transfer Data at Scale (CVPR 2025)
★ 35DeepDeblur_release. Deep Multi-scale CNN for Dynamic Scene Deblurring
★ 734prism-bench. This is the official repository for the paper "FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark"
★ 131RealBlur. Repository for Real-World Blur Dataset for Learning and Benchmarking Deblurring Algorithms
★ 202ZHO-nano-banana-Creation. 我的 nano-banana 创意玩法大合集! 持续更新中!
★ 3.7kTeaCache. Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
★ 1.4kLightX2V-Qwen-Image-Lightning. Qwen-Image-Lightning: Speed up Qwen-Image model with distillation
★ 1.3kDreamOmni2. This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
★ 2kOpenUni. Python
★ 189HunyuanImage-3.0. HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
★ 3.2kms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kAwesome-Nano-Banana-images. A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the community's development of image generation and unified models(click to website to see our blog)
★ 23kSkyReels-A2. SkyReels-A2: Compose anything in video diffusion transformers
★ 713UMO. [CVPR 2026] 🔥🔥 Official Repo of UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
★ 190reconstruction-alignment. [ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Learning.
★ 411ross. [ICLR'25] Reconstructive Visual Instruction Tuning
★ 135TinyBeauty. Python
★ 70SparkUI-Parser.
★ 23Pref-GRPO. Official implementation of Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
★ 276Lumina-Image-2.0. Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
★ 1kawesome-nano-banana. Awesome curated collection of images and prompts generated by gemini-2.5-flash-image (aka Nano Banana) state-of-the-art image generation and editing model. Explore AI generated visuals created with Gemini, showcasing Google’s advanced image generation capabilities.
★ 8.8kEcho-4o. Echo-4o: Harnessing Proprietary Models’ Synthetic Images for Improved Image Generation
★ 506diffusion-4k. [CVPR 2025] Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
★ 364ICTHP. Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment [ICCV 2025] - Official implementation
★ 45NextStep-1. [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s Multimodal Intelligence team.
★ 693UniPic. Open-source SOTA multi-image editing model
★ 871X2Edit. AAAI2026 X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
★ 97I2EBench. [NeurIPS'24] I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
★ 35echomimic_v3. [AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
★ 999bd3lms. [ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
★ 1kHPSv3. Official implementation of HPSv3: Towards Wide-Spectrum Human Preference Score (ICCV2025)
★ 331PixelHacker. PixelHacker: Image Inpainting with Structural and Semantic Consistency
★ 584Qwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kMixGRPO. [ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
★ 1.2kX-Omni. Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).
★ 427GPT-Image-Edit. GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
★ 243GLM-V. GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 2.4kDressCode. DressCode: Autoregressively Sewing and Generating Garments from Text Guidance.
★ 293GP-VTON. Official Implementation for CVPR2023 paper "GP-VTON: Towards General Purpose Virtual Try-on via Collaborative Local-Flow Global-Parsing Learning"
★ 507TryOn-Adapter. Official repository for our paper, "TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On"
★ 51Magic-TryOn. MagicTryOn is a video virtual try-on framework based on a large-scale video diffusion Transformer.
★ 568paperswithcode-data. The full dataset behind paperswithcode.com
★ 929AnyDressing. [CVPR 2025] Official implementation of "AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models"
★ 330IMAGPose. [NeurIPS 2024] 🕺IMAGPose🕺: A Unified Conditional Framework for Pose-Guided Person Generation. IMAGPose enables versatile pose-guided image generation with high detail fidelity, pose alignment, and cross-view consistency, overcoming limitations of existing methods.
★ 188ai-deadlines. :alarm_clock: AI conference deadline countdowns
★ 6kColorFlow. The official implementation of paper "ColorFlow: Retrieval-Augmented Image Sequence Colorization". ColorFlow:基于检索增强的图像序列上色
★ 462DDColor. [ICCV 2023] DDColor: Towards Photo-Realistic Image Colorization via Dual Decoders
★ 1.5kUniCTokens. A framework for unified personalized model, achieving mutual enhancement between personalized understanding and generation. Demonstrating the potential of cross-task information transfer in personalized scenario, paving the way for the development of general unified models.
★ 131Ling. Ling is a MoE LLM provided and open-sourced by InclusionAI.
★ 264RetouchGPT. This is the official code of AAAI 2025: RetouchGPT: LLM-based Interactive High-Fidelity Face Retouching via Imperfection Prompting.
★ 125CRHD-3K. The first high-definition cloth retouching dataset CRHD-3K.
★ 104GUI-G2. [AAAI 2026] GUI-G²: Gaussian Reward Modeling for GUI Grounding
★ 311PixelFreeEffects. ✨ 望图 PixelFree 美颜SDK - 全平台美颜特效引擎(iOS/Android/harmonyOS/windows/macOS/linux)| 直播/短视频/相机/照片处理 | 高性能+轻量级
★ 833FaceScore. Official repo for 【FaceScore: Benchmarking and Enhancing Face Quality in Human Generation】
★ 84deepface. A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python
★ 23kPPR10K. Official Implementation and Dataset of "PPR10K: A Large-Scale Portrait Photo Retouching Dataset with Human-Region Mask and Group-Level Consistency", CVPR 2021
★ 346fluxgym. Dead simple FLUX LoRA training UI with LOW VRAM support
★ 3.2kLatentSync. Taming Stable Diffusion for Lip Sync!
★ 5.9kFluxText. Implementation of "FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing"
★ 857Manifesto-against-the-Plagiarist-Yunhe-Wang. 讨贼王云鹤檄文
★ 1.1kTrue-Story-of-Pangu. 诺亚盘古大模型研发背后的真正的心酸与黑暗的故事。
★ 12kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kComfyUI-FluxTrainer. Python
★ 1.2kT2ITrainer. Practice Code for text to image trainer
★ 559Ovis-U1. An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image editing within a single powerful framework.
★ 450ai-toolkit. The ultimate training toolkit for finetuning diffusion models
★ 11kvedadet. A single stage object detection toolbox based on PyTorch
★ 502ShareGPT-4o-Image. Python
★ 285LLaVA-NeXT. Python
★ 4.7kMindOmni. [NeurIPS2025] The official implementation of MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
★ 139BLIP3o. Official implementation of BLIP3o-Series
★ 1.7kOmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kAwesome-Unified-Multimodal-Models. Awesome Unified Multimodal Models
★ 1.3kRISEBench. [NIPS 2025 DB Oral] Official Repository of paper: Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
★ 155