This is your work, valued
sca-cnn.cvpr17. Image Captions Generation with Spatial and Channel-wise Attention
★ 212sp-aen.cvpr18. Zero-Shot Visual Recognition using Semantic-Preserving Adversarial Embedding Networks
★ 47Thesis-latex. my Ph.D. thesis (Zhejiang University)
★ 38WSAG. [EMNLP'22] Weakly-Supervised Temporal Article Grounding
★ 14faster-rcnn.pytorch. fork from https://github.com/jwyang/faster-rcnn.pytorch
★ 10YouwikiHow. YouwikiHow dataset for weakly-supervised article grounding
★ 5Seminar. copy from https://github.com/D-X-Y/Seminar-X
★ 5S3D_Feature_Extractors. Extract S3D Video Features
★ 2FlowCIR. [ECCV 2026] FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
★ 4AVTok. [ECCV 2026] AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
★ 8LISA. [arXiv 2026] The Pytorch Implementation of LISA
★ 6SwiftI2V. [arXiv 2026] Project page for paper "SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation"
★ 86FlowMotion. [CVPR 2026] Official PyTorch implementation of "FlowMotion: Training-Free Flow Guidance for Video Motion Transfer"
★ 6FlowComposer. [CVPR 2026] FlowComposer: Composable Flows for Compositional Zero-Shot Learning
★ 4MoKus. [ECCV 2026] MoKus: This repo is the official implementation of "MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization"
★ 12Coarse-guided-Gen. [arXiv 2026] Official PyTorch Repository for "Coarse-Guided Visual Generation via Weighted h-Transform Sampling"
★ 42Awesome-Latent-Space. A paper list of Awesome Latent Space.
★ 952AAAI26-HUG. Official Implementation for AAAI26 Oral: Heterogeneous Uncertainty-Guided Composed Image Retrieval
★ 6Awesome-Latent-CoT. This repository contains a regularly updated paper list for LLMs-reasoning-in-latent-space.
★ 367BA-solver. [ICML 2026] Bi-Anchor Interpolation Solver for Accelerating Generative Modeling; Paper link: https://arxiv.org/abs/2601.21542
★ 14DualSpeed. Fast-Slow Efficient Training for Multimodal Large Language Models via Visual Token Pruning
★ 3FlowDC. [CVPR 2026 Highlight] official implementation of the paper: "FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing"
★ 7STAMP. [CVPR 2026] STAMP: Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
★ 39DCD. [NeurIPS 2025] Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models
★ 3FlowCycle. [ECCV 2026] The Pytorch Implementation of FlowCycle
★ 13FMA. [ICLR 2026] Official implementation of the paper "Exploring Cross-Modal Flows for Few-Shot Learning".
★ 24GIR-Bench. [ICLR 2026] GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
★ 35ACC. [NeurIPS 2025] Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
★ 10NoOp. [NeurIPS 2025] The Pytorch Implementation of NoOp
★ 9ICCV25-HLFormer. [ICCV 2025] Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
★ 62FreeEvent. [ICML 2025] Official PyTorch implementation of Event-Customized Image Generation
★ 4DyME. [ICLR 2026] Empowering Small VLMs to Think with Dynamic Memorization and Exploration
★ 18CVPR25-Condenser. The code for the paper "Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning" (CVPR'25).
★ 16Relation-R1. [AAAI 2026] Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension
★ 20Awesome-DiT-Inference. 📚A curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.🎉
★ 578B2-DiffuRL. [CVPR 25] A framework named B^2-DiffuRL for RL-based diffusion model fine-tuning.
★ 57IterIS-merging. [CVPR 2025] IterIS: Iterative Inference-Solving Alignment for LoRA Merging
★ 10Diff-II. [CVPR 2025] PyTorch implementation of Diff-II
★ 29CLIPDrag. [ICLR 2025] Official code for Combining Text-based and Drag-based Editing for Precise and Flexible Image Editing.
★ 20Nautilus. [ICCV 2025] Nautilus: Locality-aware Autoencoder for Scalable Mesh Generation
★ 59DisPose. [ICLR2025] DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
★ 377CausalCache-VDM. Official implementation of our paper: "Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing" (ICML 2025)
★ 83PathWeave. Code for paper "LLMs Can Evolve Continually on Modality for X-Modal Reasoning" NeurIPS2024
★ 41Awesome-MLLM-Benchmarks.
★ 163reinforcement-learning-an-introduction. Solutions to exercises in Reinforcement Learning: An Introduction (2nd Edition).
★ 411DRL. Deep Reinforcement Learning
★ 4.7kSHERL. [ECCV2024] The code of "SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning"
★ 9RLexample. Some basic examples of playing with RL
★ 1.3kMIND_Distillation. Code and data for the paper: MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding (https://arxiv.org/pdf/2406.10701).
★ 3CoMM. [CVPR 2025 Highlight] Official repository for CoMM Dataset
★ 57Practical_RL. A course in reinforcement learning in the wild
★ 6.5kDeepMind-Advanced-Deep-Learning-and-Reinforcement-Learning. Advanced Deep Learning and Reinforcement Learning course taught at UCL in partnership with Deepmind
★ 862Hands-on-RL. https://hrl.boyuai.com/
★ 4.9kReinforcement-Learning-2nd-Edition-by-Sutton-Exercise-Solutions. Solutions of Reinforcement Learning, An Introduction
★ 2.4kDeepRL-Tutorials. Contains high quality implementations of Deep Reinforcement Learning algorithms written in PyTorch
★ 1.1kBook-Mathematical-Foundation-of-Reinforcement-Learning. This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."
★ 17kAwesome-Open-Vocabulary-Detection-and-Segmentation. Awesome OVD-OVS - A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future
★ 219reinforcement_learning_course_materials. Lecture notes, tutorial tasks including solutions as well as online videos for the reinforcement learning course hosted by Paderborn University
★ 1.2kintroRL. Intro to Reinforcement Learning (强化学习纲要)
★ 3.6kAwesome-Video-Diffusion-Models. [CSUR] A Survey on Video Diffusion Models
★ 2.3kSCHEMA. [ICLR 2024 Poster] SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
★ 20Awesome-Diffusion-Model-Based-Image-Editing-Methods. Diffusion Model-Based Image Editing: A Survey (TPAMI 2025)
★ 712Awesome-LLMs-for-Video-Understanding. 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
★ 3.3kKnowledgeEditingPapers. Must-read Papers on Knowledge Editing for Large Language Models.
★ 1.2kNSFC-LaTex. BibTeX Style
★ 1.6kAwesome-Text-to-Image. (ෆ`꒳´ෆ) A Survey on Text-to-Image Generation/Synthesis.
★ 2.4kCFA. [ICCV 2023] Compositional Feature Augmentation for Unbiased Scene Graph Generation
★ 15RECODE. [NeurIPS 2023] Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models
★ 23NICEST.
★ 1Multimodal-AND-Large-Language-Models. Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
★ 760awesome-multimodal-ml. Reading list for research topics in multimodal machine learning
★ 6.9kMr.Harm-EMNLP2023. Code for our EMNLP 2023 paper - Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models
★ 15Awesome-Open-Vocabulary. (TPAMI 2024) A Survey on Open Vocabulary Learning
★ 999pbdl-book. Welcome to the Physics-based Deep Learning Book v0.3 - the GenAI Edition
★ 1.4kUniPT. [CVPR2024] The code of "UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory"
★ 71LLMSurvey. The official GitHub page for the survey paper "A Survey of Large Language Models".
★ 12kAwesome-Foundation-Models. A curated list of foundation models for vision and language tasks
★ 1.2kACL2023_ChartT5. The official code implementation of the ACL 2023 Finding paper: Enhanced Chart Understanding in Vision and Language Task via Cross-modal Pre-training on Plot Table Pairs
★ 10Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kAwesome-Multimodal-LLM. Research Trends in LLM-guided Multimodal Learning.
★ 355IdealGPT. Official Code of IdealGPT
★ 39learning_research. 本人的科研经验
★ 14kLLM-in-Vision. Recent LLM-based CV and related works. Welcome to comment/contribute!
★ 871open-in-overleaf. Edit latex of any arxiv.org paper directly on overleaf
★ 213prompt-in-context-learning. Awesome resources for in-context learning and prompt engineering: Mastery of the LLMs such as ChatGPT, GPT-3, and FlanT5, with up-to-date and cutting-edge updates.
★ 2.2kawesome-gpt4. A curated list of prompts, tools, and resources regarding the GPT-4 language model.
★ 2.2kMM-REACT. Official repo for MM-REACT
★ 967viper. Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"
★ 1.7kTempCLR. [ICLR 2023] Temporal Alignment Representations with Contrastive Learning
★ 27Prompt-Engineering-Guide. 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
★ 77kChain-of-ThoughtsPapers. A trend starts from "Chain of Thought Prompting Elicits Reasoning in Large Language Models".
★ 2.1kNSFC-LaTex. TeX
★ 38prompts.chat. f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
★ 167kOpenVoc-VidVRD. Official code for the ICLR2023 paper Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
★ 43classify_by_description_release. Python
★ 177ICL_PaperList. Paper List for In-context Learning 🌷
★ 876Diffusion-Models-Papers-Survey-Taxonomy. Diffusion model papers, survey, and taxonomy
★ 3.4kTIger. The PyTorch implementation of the [ECCV'22 paper] Explicit Image Caption Editing
★ 9KDDAug. [ECCV2022] Rethinking Data Augmentation for Robust Visual Question Answering
★ 13YouwikiHow. YouwikiHow dataset for weakly-supervised article grounding
★ 5ChartQA. Python
★ 260WSAG. [EMNLP'22] Weakly-Supervised Temporal Article Grounding
★ 14sg-tech-list. :scroll: List of notable tech companies in Singapore
★ 613Unbiased_SGG.
★ 6Chart-to-text. OpenEdge ABL
★ 128FCT. Code for CVPR 2022 Oral paper: 'Few-Shot Object Detection with Fully Cross-Transformer'
★ 93awesome_lists. Awesome Lists for Tenure-Track Assistant Professors and PhD students. (助理教授/博士生生存指南)
★ 1.6kNICEST. An extension of the CVPR paper (The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph Generation)
★ 5DCNet. [ACM MM 22] Correspondence Matters for Video Referring Expression Comprehension
★ 15ECE. [ECCV'22 Poster] Explicit Image Caption Editing
★ 22Unicorn. [ECCV'22 Oral] Towards Grand Unification of Object Tracking
★ 952WS-SGG. Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph Generation, MM 2022
★ 11TransDIC. [ACM MM 2022] Rethinking the Reference-based Distinctive Image Captioning
★ 4PEVL. Source code for EMNLP 2022 paper “PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models”
★ 49Conference-Acceptance-Rate. Acceptance rates for the major AI conferences
★ 4.8kAwesome_Prompting_Papers_in_Computer_Vision. A curated list of prompt-based paper in computer vision and vision-language learning.
★ 928HumanSystemOptimization. 健康学习到150岁 - 人体系统调优不完全指南
★ 22kNICE. [CVPR'2022 Oral] The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph Generation
★ 32Awesome-CLIP. Awesome list for research on CLIP (Contrastive Language-Image Pre-Training).
★ 1.2kPrompt-align. [ICCV 2023] Prompt-aligned Gradient for Prompt Tuning
★ 170awesome-tips.
★ 4.7kcpl. CPL: Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal Learning
★ 65pyGAT. Pytorch implementation of the Graph Attention Network model by Veličković et. al (2017, https://arxiv.org/abs/1710.10903)
★ 3.1kVidSGG-BIG. Pytorch implementation of our paper Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs, which is accepted by CVPR2022
★ 47moment_detr. [NeurIPS 2021] Moment-DETR code and QVHighlights dataset
★ 349decord. An efficient video loader for deep learning with smart shuffling that's super easy to digest
★ 2.5kMIL-NCE_HowTo100M. PyTorch GPU distributed training code for MIL-NCE HowTo100M
★ 221S3D_HowTo100M. S3D Text-Video model trained on HowTo100M using MIL-NCE
★ 200acav100m. ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning. In ICCV, 2021.
★ 64video_feature_extractor. Easy to use video deep features extractor
★ 322WSTAN. Python
★ 16Awesome-Temporal-Language-Grounding-in-Videos. A curated list of grounding natural language in video and related area. :-)
★ 105VSR-guided-CIC. Human-like Controllable Image Captioning with Verb-specific Semantic Roles.
★ 36GLIP. Grounded Language-Image Pre-training
★ 2.6kSituFormer. [AAAI 2022] Official implementation of the paper Rethinking the Two-Stage Framework for Grounded Situation Recognition, AAAI 2022.
★ 13awesome-semi-supervised-learning. 😎 An up-to-date & curated list of awesome semi-supervised learning papers, methods & resources.
★ 1.9kal-folio. A beautiful, simple, clean, and responsive Jekyll theme for academics
★ 16kccf-deadlines. ⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
★ 9.2kawesome-text-summarization. A curated list of resources dedicated to text summarization
★ 1.5kPSVL. Code for the paper "Zero-shot Natural Language Video Localization" (ICCV2021, Oral).
★ 48faiss. A library for efficient similarity search and clustering of dense vectors.
★ 41kReLoCLNet. Video Corpus Moment Retrieval with Contrastive Learning (SIGIR 2021)
★ 58LPNet. Python
★ 13VisualDS. Python
★ 24CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kcoot-videotext. COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
★ 291VidVRD-tracklets. Video Visual Relation Detection (VidVRD) tracklets generation. also for ACM MM Visual Relation Understanding Grand Challenge
★ 40Transformer-in-Vision. Recent Transformer-based CV and related works.
★ 1.3kCrossFormer. The official code for the paper: https://openreview.net/forum?id=_PHymLIxuI
★ 402swig. Situation With Groundings (SWiG) dataset and Joint Situation Localizer (JSL)
★ 71Awesome-Temporal-Action-Localization. A curated list of temporal action localization/detection and related area (e.g. temporal action proposal) resources.
★ 588cfvqa. [CVPR 2021] Counterfactual VQA: A Cause-Effect Look at Language Bias
★ 136grounding_changing_distribution.
★ 36ref-nms. Official codebase for "Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding"
★ 22awesome-visual-question-answering. A curated list of Visual Question Answering(VQA)(Image/Video Question Answering),Visual Question Generation ,Visual Dialog ,Visual Commonsense Reasoning and related area.
★ 672simple-resume-cv. Template for a simple resume or curriculum vitae (CV), in XeLaTeX.
★ 555Causal_Reading_Group. We will keep updating the paper list about machine learning + causal theory. We also internally discuss related papers between NExT++ (NUS) and LDS (USTC) by week.
★ 515hierarchical_fashion_graph_network. Hierarchical Fashion Graph Network for Personalized Outfit Recommendation, SIGIR 2020
★ 91CSS-VQA. Counterfactual Samples Synthesizing for Robust VQA
★ 78awesome-grounding. awesome grounding: A curated list of research papers in visual grounding
★ 1.1kresume. An elegant \LaTeX\ résumé template. 大陆镜像 https://gods.coding.net/p/resume/git
★ 11kSceneGraphParser. A python toolkit for parsing captions (in natural language) into scene graphs (as symbolic representations).
★ 595bottom-up-attention-vqa. An efficient PyTorch implementation of the winning entry of the 2017 VQA Challenge.
★ 768tensorboardX. tensorboard for pytorch (and chainer, mxnet, numpy, ...)
★ 8ktutorials. PyTorch tutorials.
★ 9.3kmoviepy. Video editing with Python
★ 15kuse_vim_as_ide. use vim as IDE
★ 9.2k