This is your work, valued
Comfyui-QwenEditUtils. Python
★ 830T2ITrainer. Practice Code for text to image trainer
★ 559Comfyui-In-Context-Lora-Utils. Python
★ 245ComfyUI-EditUtils. a set of utils for comfyui edit model
★ 184Comfyui-LatentUtils. a set of utils for comfyui latent operation
★ 100ComfyUI-LoaderUtils. Python
★ 86ComfyUI-Watermark-Detection. watermark detection utils
★ 51Comfyui-LoraUtils. a set of utils for comfyui lora operation
★ 32Kontext-Bench-Analysis. Python
★ 29ConsistencyDecoderNode. Python
★ 17Comfyui-Kolors-Utils. Utils for kolors
★ 17Comfyui-DiffusersUtils. Python
★ 14Comfyui-Condition-Utils. A tool to manage condition
★ 13the-super-tiny-compiler. :snowman: Possibly the smallest compiler ever
★ 9ComfyUI-VAE-Utils. Python
★ 6Comfyui-ThinkRemover. Remove content inside <think/>
★ 5CaptionFlow. Python
★ 5ComfyUIJasonNode. Some Custom Node I made for comfyui
★ 4kosmos-auto-captions. Python
★ 3AISlide. Share space.bilibili.com/443010 video doc
★ 3CaptionFusionator. PowerShell
★ 3DatasetManagement. DatasetManagement is builded for easily manage image and captions.
★ 2Comfyui-LTXCondReplace. Replace latent Cond
★ 2ComfyUI_mistral_api. A ComfyUI custom node that integrates Mistral AI's Pixtral Large vision model, enabling powerful multimodal AI capabilities within ComfyUI. Pixtral Large is a 124B parameter model (123B decoder + 1B vision encoder) that can analyze up to 30 high-resolution images simultaneously.
★ 1comfy-model-tools. Utility scripts for packaging models for ComfyUI.
★ 65Nomi. Open-source, local-first desktop app for AI video creation: write a script → generate images & video → edit on a timeline → export. Bring your own model & API key — everything runs on your machine. Built with Electron + React.
★ 374GenClaw. GenClaw: Code-Driven Agentic Image Generation
★ 298SCOPE. Python
★ 32PiD. PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
★ 997oh-my-pi. ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
★ 21kskills. Skills for Real Engineers. Straight from my .agents directory.
★ 194kLakonLab. Official implementation of AsymFlow, pi-Flow, GMFlow
★ 457D-OPSD. Official Repo of "D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models"
★ 291hello-agents. 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
★ 69kHiDream-O1-Image. Python
★ 1.5kComfyUI_RH_OpenAPI. This is a ComfyUI plugin for https://www.runninghub.cn/call-api/standard-api
★ 125FD-Loss. Python
★ 549GMNet. [ICLR2025] Learning Gain Map for Inverse Tone Mapping
★ 34openexr-images. Collection of images associated with the OpenEXR distribution
★ 132SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4kReal-HDRV. CVPR 2024: Official repository of 'Towards Real-World HDR Video Reconstruction: A Large-Scale Benchmark Dataset and A Two-Stage Alignment Network'
★ 33agentscope. Build and run agents you can see, understand and trust.
★ 28kComfyUI_Gear. VFX-oriented nodes for ComfyUI: HDR LogC3 decode + interactive ACEScct color grading panel (color wheels, scopes, A|B compare, batch scrubber) powered by exr-viewer.
★ 35LumiPic. Single-Image HDR Reconstruction via LogC3-Encoded Diffusion Transformer LoRA
★ 41houtini-lm. MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.
★ 102ComfyUI-LCS. Python
★ 153LCS. [ICML 2026] The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
★ 27QwenPaw. Your Personal AI Assistant; easy to install, deploy on your own machine or on the cloud; supports multiple chat apps with easily extensible capabilities.
★ 30kUAE. Official repo for UAE [ECCV 2026]
★ 208mlx-drifting-model. Generative Modeling via Drifting in MLX
★ 43comfyui-node-organizer. Automatically align and organize nodes in your workflow
★ 55becominglit-dataset. Official download kit of the BecomingLit dataset from the NeurIPS 2025 paper: 'BecomingLit: Relightable Gaussian Avatars with Hybrid Neural Shading'
★ 20GLM-Image. GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
★ 1kComfyUI-ReservedVRAM. A simple node that can dynamically adjust the reserved memory of a workflow in real-time, used to avoid the utilization of shared memory.
★ 389vision-pt. Python
★ 5SFD. [CVPR 2026] Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
★ 153TwinFlow. [ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻
★ 537LongCat-Image. Python
★ 713PaDT. [ICLR 2026] Official implementation of "Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs"
★ 163iMontage. iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
★ 188jit_toy_example. Unofficial implementation of the toy example in JiT https://arxiv.org/abs/2511.13720
★ 65weave. Official Repo for Paper <WEAVE: Unleashing and Benchmarking the Interleaved Cross-modality Comprehension and Generation>
★ 42ComfyUI-piFlow. ComfyUI Nodes for AsymFlow and pi-Flow
★ 186toon. 🎒 Token-Oriented Object Notation (TOON) – compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.
★ 25kComfyUI-VAE-Utils. Python
★ 219ChronoEdit. [ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
★ 701EPG. [ICLR2026] There is No VAE: End-To-End Pixel-Space Generative Modeling Via Self-Supervised Pre-Training
★ 152Comfyui-QwenEditUtils. Python
★ 830VTC-Bench. [ACL2026 Main] Data & Code of "Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods"
★ 35Rex-Omni. [CVPR2026] Detect Anything via Next Point Prediction
★ 1.5kMuon. Muon is an optimizer for hidden layers in neural networks
★ 2.7kmammothmoda. Python
★ 333DeltaFlow. [NeurIPS'25 Spotlight] DeltaFlow: An Efficient Multi-frame Scene Flow Estimation Method
★ 44Cold-Diffusion-Models. Official implementation of Cold-Diffusion for different transformations in pytorch.
★ 1.1kDiffusionNFT. [ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
★ 994Qwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kRobusTok. Image Tokenizer Needs Post-Training
★ 24DIM. [ICLR 2026] Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
★ 28accelerate. 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
★ 9.8kimage-gs. Official implementation of SIGGRAPH 2025 paper "Image-GS: Content-Adaptive Image Representation via 2D Gaussians"
★ 455DepthAnythingAC. Official code for the paper: Depth Anything At Any Condition
★ 344ControlNet. Let us control diffusion models!
★ 34kCRHD-3K. The first high-definition cloth retouching dataset CRHD-3K.
★ 104gemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kSiT. Official PyTorch Implementation of "SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers"
★ 1.2kMiniMax-Remover. This is the official implementation of our paper: "MiniMax-Remover: Taming Bad Noise Helps Video Object Removal"
★ 589OmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kLinGen. Python
★ 30asuka-misato. [CVPR 2025 Highlight] Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
★ 79MeanFlow. PyTorch implementation of MeanFlow & iMF (one-step generative modeling).
★ 1.2kYiXianMemo-python. YiXianMemo is a cards recording tool designed for Yi Xian: The Cultivation Card Game.
★ 11ObjectClear. [CVPR'26] ObjectClear: Precise Object and Effect Removal with Adaptive Target-Aware Attention
★ 608nanoDiT. Just another reasonably minimal repo for class-conditional training of pixel-space diffusion transformers.
★ 157ComfyUI-BrushNet. ComfyUI BrushNet nodes
★ 947Bagel. Open-source unified multimodal model
★ 6.1kRORem. [CVPR2025] RORem: Training a Robust Object Remover with Human-in-the-Loop
★ 70DanceGRPO. An official implementation of DanceGRPO: Unleashing GRPO on Visual Generation
★ 1.6kZenCtrl. In-context subject-driven image generation while preserving foreground fidelity
★ 351Omnieraser. Python
★ 135T2I-R1. [NeurIPS 2025] T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
★ 433Flow-Inference-Time-Scaling. [NeurIPS 2025] Official code for Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
★ 75ICEdit. [NeurIPS 2025] Image editing is worth a single LoRA! 0.1% training data for fantastic image editing! Surpasses GPT-4o in ID persistence~ MoE ckpt released! Only 4GB VRAM is enough to run!
★ 2.1kxAR. This repository includes the official implementation of our paper "Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation"
★ 251Step1X-Edit. A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.
★ 2.2kComfyUI-FramePackWrapper. Python
★ 98perception_models. State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
★ 2.3ksmart-card-workshop. An AI-powered content conversion tool that transforms text, web content, or HTML code into beautifully designed card images.一款基于AI的内容转换工具,可以将文本、网页内容或HTML代码转换为精美的卡片图像。
★ 34REPA-E. [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers
★ 511ZoomEye. [EMNLP-2025 Oral] ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
★ 91REPA. [ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
★ 1.7kFramePack. Lets make video diffusion practical!
★ 17kCobra. [SIGGRAPH 2025] Official code of the paper "Cobra: Efficient Line Art COlorization with BRoAder References". Cobra:利用更广泛参考图实现高效线稿上色
★ 256steplaw. Python
★ 235HiDream-E1. Python
★ 790DDT. [CVPR 2026] DDT: Decoupled Diffusion Transformer
★ 407VisualCloze. [ICCV 2025] VisualCloze: A universal image generation framework that can support a wide range of in-domain tasks and generalize to unseen ones. (🔥 🔥 🔥 Merged into offical pipelines of diffusers.)
★ 283CFG-Zero-star. Official repo for CFG-Zero*
★ 715RADIO. Official repository for "AM-RADIO: Reduce All Domains Into One"
★ 1.9kMemeMeow. 智能检索张维为表情包
★ 1.3kHiDream-I1. Python
★ 2.5kPixWizard. [ICLR2025] A versatile image-to-image visual assistant, designed for image generation, manipulation, and translation based on free-from user instructions.
★ 211alphabet-dataset. Synthetic Alphabet Dataset
★ 19imm. Official implementation of Inductive Moment Matching
★ 585Edit-Transfer. Official code of "Edit Transfer: Learning Image Editing via Vision In-Context Relations"
★ 89OmniPaint. [ICCV 25] OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting
★ 328tiled_ksampler. Python
★ 103thera. [TMLR 2025] Thera: Aliasing-Free Arbitrary-Scale Super-Resolution with Neural Heat Fields
★ 865IterComp. [ICLR 2025] IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation
★ 203tread. Python
★ 182DiT. Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
★ 8.7kfractalgen. PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
★ 1.2kml-gbc. Python
★ 103mdlm. [NeurIPS 2024] Simple and Effective Masked Diffusion Language Model
★ 704OmniParser. A simple screen parsing tool towards pure vision based GUI agent
★ 25kWatermark-Removal-Pytorch. 🔥 CNN for Watermark Removal using Deep Image Prior with Pytorch 🔥.
★ 1.1khanzi-writer-data. The data used by Hanzi Writer
★ 733DiT-Extrapolation. Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (ICML 2025) , UltraViCo (ICLR 2026) and UltraImage
★ 820DeepEP. DeepEP: an efficient expert-parallel communication library
★ 9.9keqvae. [ICML'25] EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling.
★ 182FlowEdit. Official implementation of the paper: "FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models"
★ 1kDiffusion-Sharpening. Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening
★ 72MaterialFusion. A framework for high-quality material transfer that allows users to adjust the degree of material application.
★ 63RAS. An open-source implementation of Regional Adaptive Sampling (RAS), a novel diffusion model sampling strategy that introduces regional variability in sampling steps
★ 153MakeAnything. Official code of "MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation"
★ 212FlashVideo. [AAAI-2026]FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
★ 460syncd. SynCD: Generating Multi-Image Synthetic Data for Text-to-Image Customization (ICCV 2025)
★ 155nodezator. A generalist Python node editor
★ 2.8kACE_plus. Python
★ 1.4kTinyZero. Minimal reproduction of DeepSeek R1-Zero
★ 13kComfyUI-ColorshiftColor. Python
★ 53DiffusionLight. [CVPR 2024] code release for "DiffusionLight: Light Probes for Free by Painting a Chrome Ball"
★ 712LyCORIS. Lora beYond Conventional methods, Other Rank adaptation Implementations for Stable diffusion.
★ 2.5kFireFlow-Fast-Inversion-of-Rectified-Flow-for-Image-Semantic-Editing. [ICML2025] An 8-step inversion and 8-step editing process works effectively with the FLUX-dev model. (3x speedup with results that are comparable or even superior to baseline methods)
★ 295latent-diffusion. High-Resolution Image Synthesis with Latent Diffusion Models
★ 14kdifferential-diffusion. Python
★ 452genesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kOneDiffusion. Official implementation of OneDiffusion paper (CVPR 2025)
★ 662Golden-Noise-for-Diffusion-Models. [ICCV2025] The code of our work "Golden Noise for Diffusion Models: A Learning Framework".
★ 197ComfyUI-NPNet. https://github.com/xie-lab-ml/Golden-Noise-for-Diffusion-Models for ComfyUI
★ 18ComfyUI_mistral_api. A ComfyUI custom node that integrates Mistral AI's Pixtral Large vision model, enabling powerful multimodal AI capabilities within ComfyUI. Pixtral Large is a 124B parameter model (123B decoder + 1B vision encoder) that can analyze up to 30 high-resolution images simultaneously.
★ 1ComfyUI_AdvancedRefluxControl. Python
★ 652Comfyui-In-Context-Lora-Utils. Python
★ 245qwen2vl-flux. Python
★ 572OminiControl. [ICCV 2025 Highlight] OminiControl: Minimal and Universal Control for Diffusion Transformer
★ 1.9kSubjects200K. Subjects200K dataset
★ 132ComfyUI-APQNodes. useful custom nodes for ComfyUI
★ 110Comfyui_Object_Migration. This is a study aim to transfer the single concept by using DIT model self-attention capablity
★ 784In-Context-LoRA. Official repository of In-Context LoRA for Diffusion Transformers
★ 2.1kHiDiffusion. [ECCV 2024] HiDiffusion: Increases the resolution and speed of your diffusion model by only adding a single line of code!
★ 840FiT. [ICML 2024 Spotlight] FiT: Flexible Vision Transformer for Diffusion Model
★ 434ComfyUI-Fluxtapoz. Nodes for image juxtaposition for Flux in ComfyUI
★ 1.4ktext-clustering. Easily embed, cluster and semantically label text datasets
★ 609T2ITrainer. Practice Code for text to image trainer
★ 559OmniGen. OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
★ 4.3kSana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6kJanus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kEagle. Eagle: Frontier Vision-Language Models with Data-Centric Strategies
★ 3.3kMPS. Python
★ 206SPRIGHT. [ECCV 2024] Official PyTorch implementation of "Getting it Right: Improving Spatial Consistency in Text-to-Image Models"
★ 106flux. Official inference repo for FLUX.1 models
★ 26krectified-flow. 从零手搓Flow Matching(Rectified Flow)
★ 637CoMat. [NeurIPS 2024] 💫CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
★ 169bilibili-api.
★ 4.2kViPer. Python
★ 71data-juicer. Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.8kComfyUI-Kolors-MZ. Kolors的ComfyUI原生采样器实现(Kolors ComfyUI Native Sampler Implementation)
★ 581ComfyUI-KwaiKolorsWrapper. Diffusers wrapper to run Kwai-Kolors model
★ 594LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kKolors. Kolors Team
★ 4.6kPhiCookBook. This is a Phi Family of SLMs book for getting started with Phi Models. Phi a family of open sourced AI models developed by Microsoft. Phi models are the most capable and cost-effective small language models (SLMs) available, outperforming models of the same size and next size up across a variety of language, reasoning, coding, and math benchmarks
★ 3.8kHunyuanDiT. Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
★ 4.3kTroL. [EMNLP 2024] Official PyTorch implementation code for realizing the technical part of Traversal of Layers (TroL) presenting new propagation operation to get super vision language performances.
★ 99Lumina-T2X. Lumina-T2X is a unified framework for Text to Any Modality Generation
★ 2.2kComfyUI-LuminaWrapper. Python
★ 196SD-Latent-Interposer. A small neural network to provide interoperability between the latents generated by the different Stable Diffusion models.
★ 327Inf-DiT. Official implementation of Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
★ 448modern_ai_for_beginners. modern AI for beginners
★ 242Long-CLIP. Scripts for use with LongCLIP, including fine-tuning Long-CLIP
★ 63Perturbed-Attention-Guidance. Official implementation of "Perturbed-Attention Guidance"
★ 328InternLM-XComposer. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2.9kcompel. A prompting enhancement library for transformers-type text embedding systems
★ 605PixArt-sigma. PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
★ 1.9kschedule_free. Schedule-Free Optimization in PyTorch
★ 2.3kLaVi-Bridge. [ECCV 2024] Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
★ 300ComfyUI-layerdiffuse. Layer Diffuse custom nodes
★ 1.8ksd-forge-layerdiffuse. [WIP] Layer Diffusion for WebUI (via Forge)
★ 4.1kNeural-Network-Diffusion. We introduce a novel approach for parameter generation, named neural network parameter diffusion (p-diff), which employs a standard latent diffusion model to synthesize a new set of parameters
★ 886imgutils. A convenient and user-friendly anime-style image data processing library that integrates various advanced anime-style image processing models
★ 405waifuc. Efficient Train Data Collector for Anime Waifu
★ 404PickScore. Python
★ 601PIA. [CVPR 2024] PIA, your Personalized Image Animator. Animate your images by text prompt, combing with Dreambooth, achieving stunning videos. PIA,你的个性化图像动画生成器,利用文本提示将图像变为奇妙的动画
★ 975sliders. Concept Sliders for Precise Control of Diffusion Models
★ 1.1kComfyUI_ExtraModels. Support for miscellaneous image models. Currently supports: DiT, PixArt, HunYuanDiT, MiaoBi, and a few VAEs.
★ 537ComfyUI-EasyNode. JavaScript
★ 68onediff. OneDiff: An out-of-the-box acceleration library for diffusion models.
★ 2kPoolNet. Code for our CVPR 2019 paper "A Simple Pooling-Based Design for Real-Time Salient Object Detection"
★ 59consistencydecoder. Consistency Distilled Diff VAE
★ 2.2kAwesome-Aesthetic-Evaluation-and-Cropping.
★ 409PixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kgenerative-models. Generative Models by Stability AI
★ 27kpdfmake. Client/server side PDF printing in pure JavaScript
★ 12kvue-content-loader. SVG component to create placeholder loading, like Facebook cards loading.
★ 3kvue-api-query. 💎 Elegant and simple way to build requests for REST API
★ 1.7k