This is your work, valued
PhD student at Tsinghua University
DreamLight. Python
★ 395UniLSeg. [CVPR 2024] Official implementation of "Universal Segmentation at Arbitrary Granularity with Language Instruction"
★ 283QDMN. [ECCV2022] Learning Quality-aware Dynamic Memory for Video Object Segmentation
★ 124SCAN. [CVPR 2024] The repository contains the official implementation of "Open-Vocabulary Segmentation with Semantic-Assisted Calibration"
★ 77Awesome-Unified-Understanding-and-Generation.
★ 52GSFM. [ECCV2022] Global Spectral Filter Memory Network for Video Object Segmentation
★ 42MRVS_SOC. Python
★ 9GKC. [ICCV2023] The repository contains the implementation of "Global Knowledge Calibration for Fast Open-Vocabulary Segmentation"
★ 8DAVIS-evaluation. This repository (Python3) is used to easily evaluate DAVIS2016 and DAVIS2017 val set
★ 6yongliu20.github.io. HTML
★ 2IBFlow. Python
★ 12AutoResearchClaw. Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
★ 14kAgentsMeetRL. Awesome List for Agentic RL
★ 1.7kMeta-CoT. [CVPR 2026] Official code of the paper "Meta-CoT: Enhancing Granularity and Generalization in Image Editing"
★ 79MOVA. MOVA: Towards Scalable and Synchronized Video–Audio Generation
★ 1.1kDiffusionNFT. [ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
★ 998JoyAI-Image. JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
★ 2.2kbest-skills. 通用高质量 Skills 合集🔥
★ 2.4kAwesome-RL-for-Multimodal-Foundation-Models. 📖 This is a repository for organizing papers, codes and other resources related to Visual Reinforcement Learning.
★ 452Awesome-RL-for-Video-Generation. A curated list of papers on reinforcement learning for video generation
★ 575TempFlow-GRPO. [ICLR 26] TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.
★ 480dingo. Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
★ 730ml-l3m. Large multi-modal models (L3M) pre-training.
★ 229img2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4.4kSRPO. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
★ 1.3kAwesome-Nano-Banana-images. A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the community's development of image generation and unified models(click to website to see our blog)
★ 23kMini-o3. Official Code for "Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search"
★ 423USP. [ICCV25] USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
★ 96AICGSecEval. A.S.E (AICGSecEval) is a repository-level AI-generated code security evaluation benchmark developed by Tencent Wukong Code Security Team.
★ 647VisRAG. Parsing-free RAG supported by VLMs
★ 975VLM-R1. Solve Visual Understanding with Reinforced VLMs
★ 6kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kTokLIP. TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
★ 236Academic-LaTeX-Writing-Submission-Checklist-. This checklist is designed to help you systematically prepare and polish academic papers for top conferences and journals (e.g., ICML, NeurIPS, CVPR). It incorporates widely recommended best practices, formatting standards, and common reviewer expectations.
★ 226Awesome-Unified-Understanding-and-Generation.
★ 52diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kOmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1ktrl. Train transformer language models with reinforcement learning.
★ 19kDreamLight. Python
★ 395Harmon. [ICCV2025]Code Release of Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
★ 192Awesome-Controllable-T2I-Diffusion-Models. A collection of resources on controllable generation with text-to-image diffusion models.
★ 1.1kUniVG-R1. UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
★ 167FlexiAct. [SIGGRAPH 2025] Official code of the paper "FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios"
★ 342VISA. [ECCV24] VISA: Reasoning Video Object Segmentation via Large Language Model
★ 214vlarl. Single-file implementation to advance vision-language-action (VLA) models with reinforcement learning.
★ 447Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kSa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.7kboxmot. BoxMOT: Pluggable Python and C++ SOTA multi-object tracking modules with support for axis-aligned and oriented bounding boxes
★ 8.3kCastDet. [ECCV'24/IJCV'26] Code repo for "Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher Learning"
★ 84awesome-object-detection-in-aerial-images. A curated list of awesome resources for generic object detection in aerial images.
★ 158HyperSeg. [CVPR2025] Project for "HyperSeg: Towards Universal Visual Segmentation with Large Language Model".
★ 183AnyBimanual. [ICCV2025] AnyBimanual: Transfering Unimanual Policy for General Bimanual Manipulation
★ 102Momentum-GS. [ICCV 2025] Code for Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction
★ 173PowerPaint. [ECCV 2024] PowerPaint, a versatile image inpainting model that supports text-guided object inpainting, object removal, image outpainting and shape-guided object inpainting with only a single model. 一个高质量多功能的图像修补模型,可以同时支持插入物体、移除物体、图像扩展、形状可控的物体生成,只需要一个模型
★ 1.1kSC-CLIP. [TIP 2025] Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
★ 73QVLM. [NeurIPS'24]Efficient and accurate memory saving method towards W4A4 large multi-modal models.
★ 102Flash-VStream. This is the official implementation of ICCV 2025 "Flash-VStream: Efficient Real-Time Understanding for Long Video Streams"
★ 287SEED-Voken. SEED-Voken: A Series of Powerful Visual Tokenizers
★ 1kUniRef. [ICCV2023] Segment Every Reference Object in Spatial and Temporal Spaces
★ 238VIPOSeg-Benchmark. The benchmark for "Video Object Segmentation in Panoptic Wild Scenes".
★ 12VideoCrafter. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
★ 5.1kSegLossOdyssey. A collection of loss functions for medical image segmentation
★ 4kPSALM. [ECCV2024] This is an official implementation for "PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model"
★ 270ffhq-dataset. Flickr-Faces-HQ Dataset (FFHQ)
★ 4.2kDiffiT. [ECCV 2024] Official Repository for DiffiT: Diffusion Vision Transformers for Image Generation
★ 517IC-Light. More relighting!
★ 8.5kCVPR-2024-Papers.
★ 1.1kVoCo-LLaMA. [CVPR'2025] VoCo-LLaMA: This repo is the official implementation of "VoCo-LLaMA: Towards Vision Compression with Large Language Models".
★ 205VisionLLM. VisionLLM Series
★ 1.2kDPMesh. The repository contains the official implementation of "DPMesh: Exploiting Diffusion Prior for Occluded Human Mesh Recovery", CVPR 2024
★ 45FlowIE. [CVPR 2024 oral]This repository contains the official implementation of "FlowIE: Efficient Image Enhancement via Rectified Flow"
★ 153FourierTransformer. The official Pytorch implementation of the paper "Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator" (ACL 2023 Findings)
★ 40CoHD. The official implementation of A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
★ 27iLLaMA. Adapting LLaMA Decoder to Vision Transformer
★ 30MotionLCM. [ ECCV 2024 ] MotionLCM: This repo is the official implementation of "MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model"
★ 462FLAG3Dv2. Python
★ 25efficient-kan. An efficient pure-PyTorch implementation of Kolmogorov-Arnold Network (KAN).
★ 4.7kMoCha-Stereo. [CVPR2024] The official implementation of "MoCha-Stereo: Motif Channel Attention Network for Stereo Matching”. & [arXiv] The official implementation of "Motif Channel Opened in a White-Box: Stereo Matching via Motif Correlation Graph"
★ 161T-Rex. [ECCV2024] API code for T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
★ 2.7kllama3. The official Meta Llama 3 GitHub site
★ 29kcycle-diffusion. [ICCV 2023] A latent space for stochastic diffusion models
★ 658MA-LMM. (2024CVPR) MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
★ 352VAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kInfEdit. [CVPR 2024] Official implementation, Inversion-Free Image Editing with Natural Language"
★ 362awesome-text-to-image-studies. A collection of awesome text-to-image generation studies.
★ 761InstructDiffusion. PyTorch implementation of InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions.
★ 445instruct-pix2pix. Python
★ 6.9kVim. [ICML 2024] Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
★ 3.9kLOGO. [CVPR 2023] LOGO: A Long-Form Video Dataset for Group Action Quality Assessment
★ 48ManiGaussian. [ECCV 2024] ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
★ 281ov-seg. This is the official PyTorch implementation of the paper Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP.
★ 758mistral-inference. Official inference library for Mistral models
★ 11kgemma. Gemma open-weight LLM library, from Google DeepMind
★ 5.6kAwesome-Open-Vocabulary. (TPAMI 2024) A Survey on Open Vocabulary Learning
★ 999TaPA. [arXiv 2023] Embodied Task Planning with Large Language Models
★ 195Awesome-Open-Vocabulary-Semantic-Segmentation. A curated publication list on open vocabulary semantic segmentation and related area (e.g. zero-shot semantic segmentation) resources..
★ 893awesome-described-object-detection. A curated list of papers and resources related to Described Object Detection, Open-Vocabulary/Open-World Object Detection and Referring Expression Comprehension. Updated frequently and pull requests welcomed.
★ 359OVIR-3D. This is the official repository for OVIR-3D: Open-Vocabulary 3D Instance Retrieval Without Training on 3D Data. (CoRL'23)
★ 114openscene. [CVPR'23] OpenScene: 3D Scene Understanding with Open Vocabularies
★ 846OV-3DET. Python
★ 99TripletAttention. The official implementation of Triplet Attention.
★ 23MCG-Blackbox. The MCG black-box attack framework published in TPAMI 2022
★ 38openmask3d. Python
★ 267GKC. [ICCV2023] The repository contains the implementation of "Global Knowledge Calibration for Fast Open-Vocabulary Segmentation"
★ 8Entity. EntitySeg Toolbox: Towards Open-World and High-Quality Image Segmentation
★ 1kDCI. Densely Captioned Images (DCI) dataset repository.
★ 197LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kUVCOM. [CVPR 2024] Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection
★ 117UniLSeg. [CVPR 2024] Official implementation of "Universal Segmentation at Arbitrary Granularity with Language Instruction"
★ 283AnyDoor. Official implementations for paper: Anydoor: zero-shot object-level image customization
★ 4.2kLVVIS. Large-Vocabulary Video Instance Segmentation dataset
★ 100SCAN. [CVPR 2024] The repository contains the official implementation of "Open-Vocabulary Segmentation with Semantic-Assisted Calibration"
★ 77LLaMA-VID. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
★ 861segment-caption-anything. [CVPR'24] The repository provides code for running inference and training for "Segment and Caption Anything" (SCA) , links for downloading the trained model checkpoints, and example notebooks / gradio demo that show how to use the model.
★ 233ZoneEval. Zone Evaluation: Revealing Spatial Bias in Object Detection (TPAMI 2024)
★ 46NeurIPS2023_SOC. [NeurIPS 2023] The official implementation of SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation
★ 33FLYP. Code for Finetune like you pretrain: Improved finetuning of zero-shot vision models
★ 106segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kAwesome-Temporal-Action-Localization. A curated list of temporal action localization/detection and related area (e.g. temporal action proposal) resources.
★ 588gpt_academic. 为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, moss等。
★ 71kRIFormer. Codes for "RIFormer: Keep Your Vision Backbone Effective But Removing Token Mixer"
★ 7tuning_playbook. A playbook for systematically maximizing the performance of deep learning models.
★ 30kPaint-by-Example. Paint by Example: Exemplar-based Image Editing with Diffusion Models
★ 1.3kMedSegDiff. Using Diffusion Models to Segment/Reconstruct Organs from Medical Images [AAAI Most influential Paper]
★ 1.4kglead. [CVPR 2023] GLeaD: Improving GANs with A Generator-Leading Task
★ 32SPI. [CVPR 2023] SPI: 3D GAN Inversion with Facial Symmetry Prior
★ 122awesome_lists. Awesome Lists for Tenure-Track Assistant Professors and PhD students. (助理教授/博士生生存指南)
★ 1.6kFengshenbang-LM. Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
★ 4.1kVNext. Next-generation Video instance recognition framework on top of Detectron2 which supports InstMove (CVPR 2023), SeqFormer(ECCV Oral), and IDOL(ECCV Oral))
★ 617QDMN. [ECCV2022] Learning Quality-aware Dynamic Memory for Video Object Segmentation
★ 124XMem. [ECCV 2022] XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model
★ 2kpadinv. [ECCV 2022] PadInv: High-fidelity GAN Inversion with Padding Space
★ 87ccf-deadlines. ⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
★ 9.2kAwesome-Referring-Image-Segmentation. :books: A collection of papers about Referring Image Segmentation.
★ 826IFA. Learning Implicit Feature Alignment Function for Semantic Segmentation, ECCV 2022
★ 70youtube-dl. Command-line program to download videos from YouTube.com and other video sites
★ 141kRPCMVOS. [AAAI22 Oral] Reliable Propagation-Correction Modulation for Video Object Segmentation
★ 78Robust-Video-Object-Segmentation. [ACM MM22] Towards Robust Video Object Segmentation with Adaptive Object Calibration, ACM Multimedia 2022
★ 50DAVIS-evaluation. This repository (Python3) is used to easily evaluate DAVIS2016 and DAVIS2017 val set
★ 6GSFM. [ECCV2022] Global Spectral Filter Memory Network for Video Object Segmentation
★ 42StyleHEAT. [ECCV 2022] StyleHEAT: A framework for high-resolution editable talking face generation
★ 656CRIS.pytorch. An official PyTorch implementation of the CRIS paper
★ 282MTTR. Python
★ 655ReferFormer. [CVPR2022] Official Implementation of ReferFormer
★ 356DenseCLIP. [CVPR 2022] DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
★ 549Mask-Propagation. [CVPR 2021] MiVOS - Mask Propagation module. Reproduced STM (and better) with training code :star2:. Semi-supervised video object segmentation evaluation.
★ 132STCN. [NeurIPS 2021] Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation
★ 568