This is your work, valued
PhD student at HKUST
animate-your-word. [ICCV'25 Best Paper Candidate] Official Implementations for Paper: Dynamic Typography: Bringing Text to Life via Video Diffusion Prior
★ 353MagicQuillV2. Official Implementations for Paper - MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
★ 153zliucz.github.io. Github Pages template for academic personal websites, forked from mmistakes/minimal-mistakes
★ 2next-forcing.
★ 7CameraAnything. Python
★ 65OPSD-V. On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
★ 480Understand-Anything. Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
★ 77klingbot-video. Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
★ 894lingbot-world-v2. Infinite Worlds with Versatile Interactions
★ 1.4klingbot-vision. Self-supervised learning for spatial perception
★ 888lingbot-va. [RSS 2026] Causal video-action world model for generalist robot control
★ 1.7kawesome-game-generation. 🕹️ Explore cutting-edge techniques in game generation
★ 84WorldDirector. Python
★ 69guizang-ppt-skill. AI-agent Skill for generating polished HTML slide decks: editorial magazine and Swiss layouts, image prompts, social covers, and a WebGL/low-power presentation runtime.
★ 23kcodegraph. Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent — fewer tokens, fewer tool calls, 100% local
★ 64konnx. Open standard for machine learning interoperability
★ 21kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kflashdreams. high-performance inference and serving library for interactive autoregressive video and world models
★ 433minWM. A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
★ 753WorldKV. Official implementation of "WorldKV: Efficient World Memory with World Retrieval and Compression"
★ 89xDiT. xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
★ 2.7kSynLayers. Python
★ 10Practical-RIFE. More practical frame interpolation approach.
★ 982Real-ESRGAN. Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
★ 36kSparkVSR. [ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
★ 692RealBasicVSR. Official repository of "Investigating Tradeoffs in Real-World Video Super-Resolution"
★ 1.1kECCV2022-RIFE. ECCV2022 - Real-Time Intermediate Flow Estimation for Video Frame Interpolation
★ 5.5kFlashVSR. [CVPR 2026] Towards Real-Time Diffusion-Based Streaming Video Super-Resolution — An efficient one-step diffusion framework for streaming VSR with locality-constrained sparse attention and a tiny conditional decoder.
★ 1.7kLTX-Video. Official repository for LTX-Video
★ 11kShotStream. [ECCV 2026] ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
★ 174LongLive. Long Video Gen Infrastructure
★ 2.5kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5klingbot-map. A feed-forward 3D foundation model for reconstructing scenes from streaming data
★ 16kLiveAvatar. [ECCV 2026 Oral] Implementation of "Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length"
★ 2.3kAwesome-Video-World-Models-with-AR-Diffusion. A Curated List of Awesome Video World Models with AR Diffusion: Covering Algorithms, Applications, and Infrastructure, Aimed at Serving as a Comprehensive Resource for Researchers, Practitioners, and Enthusiasts.
★ 682InstantStyle. InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation 🔥
★ 2kStoch3DGS. Cuda
★ 24HPSv2. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
★ 677daVinci-MagiHuman. Python
★ 2.1kseoul-world-model. Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis
★ 622MAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.7kopenskills. Universal skills loader for AI coding agents - npm i -g openskills
★ 11kStable-Video-Infinity. [ICLR 26 Oral] Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
★ 2.5kCausVid. (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
★ 1.4kCausal-Forcing. [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
★ 893lingbot-world. Advancing Open-source World Models
★ 4.3klingbot-vla. A Pragmatic VLA Foundation Model
★ 1.7klingbot-depth. Masked Depth Modeling for Spatial Perception
★ 1.5kControlAR. [ICLR 2025] ControlAR: Controllable Image Generation with Autoregressive Models
★ 328EditAR. EditAR: Unified Conditional Generation with Autoregressive Models (CVPR 2025)
★ 44AnyEdit. 【CVPR 2025 Oral】Official Repo for Paper "AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea"
★ 227GenerativePhotomontage. Python
★ 86iris.c. Flux 2 image generation model pure C inference
★ 2kTwinFlow. [ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻
★ 537T2I-Distill. [Tutorial] Few-Step Distillation for Text-to-Image Generation: A Practical Guide
★ 370claude-code. Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
★ 140kOmniPSD. Official code implementation of "OmniPSD: Layered PSD Generation with Diffusion Transformer"
★ 130AlphaVAE. Python
★ 91GLM-Image. GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
★ 1kLightX2V-Qwen-Image-Lightning. Qwen-Image-Lightning: Speed up Qwen-Image model with distillation
★ 1.3kLightCompress. [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
★ 737LightX2V. Lightweight Image Video Action Generation Inference Framework
★ 2.6kMAI-UI. Qwen-UI-Agent: Towards Next-Generation Real-World Centric Foundation GUI Agent
★ 1.9kCLD. Code release for: Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
★ 55piecewise-rectified-flow. PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator (NeurIPS 2024)
★ 538DMD2. (NeurIPS 2024 Oral 🔥) Improved Distribution Matching Distillation for Fast Image Synthesis
★ 1.4kdmdr. [ECCV 2026] Official Code of "Distribution Matching Distillation Meets Reinforcement Learning"
★ 288OneReward. Python
★ 348sd3.5. Python
★ 1.5ktrl. Train transformer language models with reinforcement learning.
★ 19krcm. rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
★ 777TurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kart-msra. [CVPR 2025] Official repo for ART:Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation
★ 374Z-Image. Python
★ 12kQwen-Image-Layered. Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
★ 2kWorldCanvas. Python
★ 147TimeLens. [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
★ 163Reward-Forcing. [CVPR 2026 Highlight] Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
★ 352MagicQuillV2. Official Implementations for Paper - MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
★ 153Awesome-World-Models. A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
★ 3.3kflux2. Official inference repo for FLUX.2 models
★ 2.6khed. code for Holistically-Nested Edge Detection
★ 1.9kChronoEdit. [ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
★ 701xformers. Hackable and optimized Transformers building blocks, supporting a composable construction.
★ 11kpico-banana-400k. Python
★ 1.8kdreambooth.
★ 1kNeuralSVG. Official implementation of NerualSVG
★ 1.4kHoloCine. [CVPR 2026 Highlight] Official Implementations for Paper - HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
★ 695Ditto. [CVPR'26 Highlight] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
★ 618ImgEdit. [NeurIPS 2025 D&B🔥] ImgEdit: A Unified Image Editing Dataset and Benchmark
★ 331DreamOmni2. This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
★ 2kEditVerse. Official repo for paper "EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning"
★ 139LucidFlux. LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer, ICLR 2026
★ 1.3kComfyUI-Easy-Use. In order to make it easier to use the ComfyUI, I have made some optimizations and integrations to some commonly used nodes.
★ 2.6kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kBiRefNet. [CAAI AIR'24] Bilateral Reference for High-Resolution Dichotomous Image Segmentation
★ 4kflux. Official inference repo for FLUX.1 models
★ 26kDiception. [NeurIPS 2025 Spotlight] A Generalist Diffusion Model for Vision Perception
★ 321CookLikeHOC. 🥢像老乡鸡🐔那样做饭。已添加2026年发布的《老乡鸡菜品溯源报告 2.0中新出现的菜品。主要部分于2024年完工,非老乡鸡官方仓库。文字来自《老乡鸡菜品溯源报告》,并做归纳、编辑与整理。CookLikeHOC.
★ 24kAwesome-Nano-Banana-images. A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the community's development of image generation and unified models(click to website to see our blog)
★ 23kDiffSynth-Studio. Enjoy the magic of Diffusion models!
★ 13kStand-In. [CVPR2026 🎉] Stand-In is a lightweight, plug-and-play framework for identity-preserving video generation.
★ 779awesome-nano-banana. Awesome curated collection of images and prompts generated by gemini-2.5-flash-image (aka Nano Banana) state-of-the-art image generation and editing model. Explore AI generated visuals created with Gemini, showcasing Google’s advanced image generation capabilities.
★ 8.8kDreamO. [SIGGRAPH Asia 2025] DreamO: A Unified Framework for Image Customization
★ 1.7kUNO. [ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
★ 1.4kUSO. [CVPR 2026] 🔥🔥 Official Repo of USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
★ 1.2kmnistNN. C++
★ 35Anymate. Python
★ 141SparseOcc. [ECCV 2024] Fully Sparse 3D Occupancy Prediction & RayIoU Evaluation Metric
★ 440Wan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kQwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kdress-code. Dress Code: High-Resolution Multi-Category Virtual Try-On. ECCV 2022
★ 658insightface. State-of-the-art 2D and 3D Face Analysis Project
★ 29klearning_research. 本人的科研经验
★ 14kawesome-virtual-try-on. A curated list of awesome research papers, projects, code, dataset, workshops etc. related to virtual try-on.
★ 3.1kToonCrafter. [SIGGRAPH Asia 2024, Journal Track] ToonCrafter: Generative Cartoon Interpolation
★ 6kDataFlow. Easy Data Preparation with latest LLMs-based Operators and Pipelines.
★ 7.2kOmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1klpd. [ICLR 2026 Oral] Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
★ 104radial-attention. [NeurIPS 2025] Radial Attention: O(nlogn) Sparse Attention with Energy Decay for Long Video Generation
★ 605semantic-segmentation-pytorch. Pytorch implementation for Semantic Segmentation/Scene Parsing on MIT ADE20K dataset
★ 5.1klama. 🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022
★ 10kJarvisArt. [NeurIPS' 2025] JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
★ 856Calligrapher. Calligrapher: Freestyle Text Image Customization
★ 297catvton-flux. Python
★ 614EasyControl. Implementation of "EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer"(ICCV2025)
★ 1.7kinformative-drawings. Unpaired line drawing generation
★ 448deepcompressor. Model Compression Toolbox for Large Language Models and Diffusion Models
★ 796OminiControl. [ICCV 2025 Highlight] OminiControl: Minimal and Universal Control for Diffusion Transformer
★ 1.9kcontrol-lora-v3. ControlLoRA Version 3: LoRA Is All You Need to Control the Spatial Information of Stable Diffusion.
★ 79T2ITrainer. Practice Code for text to image trainer
★ 559FLUX-Fill-LoRa-Training. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch and FLAX.
★ 120ollama. Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
★ 177kRORD. Python
★ 31Omnieraser. Python
★ 135DynamiCrafter. [ECCV 2024, Oral] DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
★ 3kpersonalize-anything. [AAAI 2026] Personalize Anything for Free with Diffusion Transformer
★ 362LayerDiffuse-Flux. Python
★ 246ComfyUI_frontend. Official front-end implementation of ComfyUI
★ 1.9kFlipSketch. FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations
★ 359insert-anything. Python
★ 574ICEdit. [NeurIPS 2025] Image editing is worth a single LoRA! 0.1% training data for fantastic image editing! Surpasses GPT-4o in ID persistence~ MoE ckpt released! Only 4GB VRAM is enough to run!
★ 2.1kerasedraw-data. [ECCV 2024] Code for "EraseDraw: Learning to Insert Objects by Erasing Them from Images"
★ 26DesignEdit. [AAAI2025] DesignEdit: Unify Spatial-Aware Image Editing via Training-free Inpainting with a Multi-Layered Latent Diffusion Framework
★ 369GoT. Official repository of "GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing"
★ 317Step1X-Edit. A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.
★ 2.2kInternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kAwesome-Object-Placement. A curated list of papers, code, and resources pertaining to object placement.
★ 110libcom. Image composition toolbox: everything you want to know about image composition/compositing or object/subject insertion/addition/compositing.
★ 734Object-Placement-Assessment-Dataset-OPA. The first dataset of composite images with rationality score indicating whether the object placement in a composite image is reasonable.
★ 86transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kQwen-VL-Series-Finetune. An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1.9kAttentiveEraser. Official implementation of the paper "Attentive Eraser: Unleashing Diffusion Model’s Object Removal Potential via Self-Attention Redirection Guidance" (AAAI 2025 Oral)
★ 222spice. SPICE: A Synergistic, Precise, Iterative, and Customizable Image Editing Workflow
★ 25Paint-by-Inpaint. Paint by Inpaint: Learning to Add Image Objects by Removing Them First
★ 117inst-inpaint. A novel inpainting framework that can remove objects from images based on the instructions given as text prompts.
★ 385Visual-RFT. Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’
★ 2.3kAniDoc. [CVPR'25] Official Implementations for Paper - AniDoc: Animation Creation Made Easier
★ 573nunchaku. [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
★ 3.9kComfyUI-nunchaku. ComfyUI Plugin of Nunchaku
★ 2.9kOmniSVG. [NeurIPS 2025] OmniSVG is the first family of end-to-end multimodal SVG generators that leverage pre-trained Vision-Language Models (VLMs), capable of generating complex and detailed SVGs, from simple icons to intricate anime characters.
★ 2.6kHQ-Edit. [ICLR 2025] HQ-Edit: A High-Quality and High-Coverage Dataset for General Image Editing
★ 114ComfyUI-to-Python-Extension. A powerful tool that translates ComfyUI workflows into executable Python code.
★ 2.4kGroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
★ 10kGrounded-SAM-2. Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2
★ 3.7kBrushEdit. [under review] The official implementation of paper "BrushEdit: All-In-One Image Inpainting and Editing"
★ 589LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kVLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3kQwen-VL. The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
★ 6.7kQwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kIMavatar. Official repository for CVPR 2022 paper: I M Avatar: Implicit Morphable Head Avatars from Videos
★ 664taesd. Tiny AutoEncoder for Stable Diffusion (and other image models)
★ 958vega-lite. A concise grammar of interactive graphics, built on Vega.
★ 5.4kComfyUI-IPAdapter-Flux. Python
★ 474InfEdit. [CVPR 2024] Official implementation, Inversion-Free Image Editing with Natural Language"
★ 362Wan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17klive_sketch. Python
★ 488MagicQuill. [CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System
★ 3.7kComfyUI-Fluxtapoz. Nodes for image juxtaposition for Flux in ComfyUI
★ 1.4kFLARE. Python
★ 722ACE. All-round Creator and Editor
★ 236ACE_plus. Python
★ 1.4kToolCommander. [NAACL 2025 Main] Official implementation of "From Allies to Adversaries: Manipulating LLM Tool Scheduling through Adversarial Injection".
★ 22Janus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kawesome-deepseek-integration. Integrate the DeepSeek API into popular software
★ 38kMagicBrush. [NeurIPS'23] "MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing".
★ 411unipaint. Code Implementation of "Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion Model"
★ 129PnPInversion. [ICLR2024] Official repo for paper "PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of Code"
★ 408SmartEraser. [CVPR 2025] Official implementation of the paper "SmartEraser: Remove Anything from Images using Masked-Region Guidance".
★ 207PuLID. [NeurIPS 2024] Official code for PuLID: Pure and Lightning ID Customization via Contrastive Alignment
★ 3.5kawesome-flux-ai. A curated list of awesome resources, tools, libraries, and applications related to Flux AI technology. This repository aims to be a comprehensive collection for developers, researchers, and enthusiasts interested in Flux AI.
★ 109rgthree-comfy. Making ComfyUI more comfortable!
★ 3.3kIDEA-Bench. Official repository of IDEA-Bench
★ 41In-Context-LoRA. Official repository of In-Context LoRA for Diffusion Transformers
★ 2.1kinclusiviz. Jupyter Notebook
★ 3Awesome-AI4Animation. [ICCVW 2025] This repository includes latest papers, projects and datasets on GenAI for Cel-Animation. Accepted by ICCV 2025 AISTORY Workshop.
★ 208FlowEdit. Official implementation of the paper: "FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models"
★ 1kMangaNinjia. [CVPR 2025 Highlight] Official implementation of "MangaNinja: Line Art Colorization with Precise Reference Following"
★ 737x-flux-comfyui. Python
★ 1.7ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kAwesome-Image-Composition. A curated list of papers, code and resources pertaining to image composition/compositing or object/subject insertion/addition/compositing, which aims to generate realistic composite image.
★ 1.1kAwesome-Generative-Image-Composition. A curated list of papers, code, and resources pertaining to generative image composition or object insertion.
★ 153