This is your work, valued
Associate Professor on Computer Vision and Machine Learning
SAKDN. [IEEE T-IP 2021] Semantics-aware Adaptive Knowledge Distillation for Cross-modal Action Recognition
★ 29TCGL. [IEEE T-IP 2022] TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning
★ 24CMCIR. [IEEE T-PAMI 2023] Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering
★ 20JSRDA. [IEEE T-CSVT 2019] Hierarchically Learned View-Invariant Representations for Cross-View Action Recognition
★ 14CDFAG. Transferable Feature Representation for Visible-to-Infrared Cross-Dataset Human Action Recognition (Complexity 2018)
★ 13VisionGRU. VisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis
★ 13DIVAFN. [IEEE T-IP 2020] Deep Image-to-Video Adaptation and Fusion Networks for Action Recognition
★ 7TSTDDs. [IEEE SPL 2018] Global Temporal Representation based CNNs for Infrared Action Recognition
★ 7MIR. [MIR 2022] Causal Reasoning Meets Visual Representation Learning: A Prospective Study
★ 3yangliu9208.github.io. HTML
★ 2VCD. VCD: Visual Causality Discovery for Cross-Modal Question Reasoning
★ 2VCSR. [ACM MM 2023] Visual Causal Scene Refinement for Video Question Answering
★ 2HORLN. Hybrid-Order Representation Learning for Electricity Theft Detection (IEEE T-II 2022)
★ 2BRIDGE-WA. BRIDGE-WA: Predicting Where and How the World Changes for Robotic Action
★ 9TurnCoder. A transparent, controllable local AI coding workbench. You see and edit every bubble in the context. Turn-based, multi-model, per-call optimized, mobile-friendly.一个透明、可控的本地AI编码工作台。您可以查看并编辑上下文中的每一个气泡。支持回合制、多模型、逐次调用优化,并兼容移动设备。
★ 13RoVLA. RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models
★ 12PhyAgentOS. PhyAgentOS is a self-evolving embodied AI operating system built on agentic workflows.
★ 1.3kDDP-WM. DDP-WM: Disentangled Dynamics Prediction for Efficient World Models (ICML-26)
★ 19TAVP. Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation (CVPR-26)
★ 25DART. DART: Differentiable Adaptive Region Tokenizer for Vision Foundation Models
★ 22CausalVLR. CausalVLR: A Toolbox and Benchmark for Vision-Language Causal Reasoning (多模态因果推理开源框架)
★ 1.1k3DAffordSplat. 3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians (ACM MM 25)
★ 79InfiniteWorld. Python
★ 86EXPRESS-Bench. Embodied Question Answering (EQA) benchmark and method (ICCV 2025)
★ 60CRA-GQA. The official implementation of "Cross-modal Causal Relation Alignment for Video Question Grounding. (CVPR 2025 Highlight)"
★ 52DSPNet. The official repository of [CVPR2025] DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
★ 28VisionGRU. VisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis
★ 13LH-VLN. Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method (CVPR-25)
★ 250ODMixer. Python
★ 13Embodied_AI_Paper_List. [Embodied-AI-Survey-2025] Paper List and Resource Repository for Embodied AI
★ 2.1kyangliu9208.github.io. HTML
★ 2Book-of-MLM. 《多模态大模型:新一代人工智能技术范式》在线资源
★ 1Book-of-MLM. 《多模态大模型:新一代人工智能技术范式》配套教学资源
★ 311ESL. Enhanced Soft Label for Semi-Supervised Semantic Segmentation. ICCV 2023
★ 5Veri3d. Python
★ 38CrossLoc. [CVPR'22] CrossLoc localization: a cross-modal visual representation learning method for absolute localization
★ 113SwinGNN. SwinGNN: diffusion model for graph generation
★ 87PSG-biased-annotation. Panoptic Scene Graph Biased Annotation
★ 35WaterScenes. Official repository for WaterScenes dataset
★ 186Awesome-Radar-Camera-Fusion. Radar Camera Fusion in Autonomous Driving
★ 424Awesome-Skeleton-based-Action-Recognition. A curated paper list of awesome skeleton-based action recognition.
★ 724ICL_SSL. [IEEE TNNLS 2022] An official source code for paper Interpolation-based Contrastive Learning for Few-Label Semi-Supervised Learning.
★ 40ChatGPT-MBTI. [EMNLP-2023] Official Codes for “Can ChatGPT Assess Human Personalities? A General Evaluation Framework”
★ 100VCD. VCD: Visual Causality Discovery for Cross-Modal Question Reasoning
★ 2DealMVC. [ACM MM 2023] An official source code for paper "DealMVC: Dual Contrastive Calibration for Multi-view Clustering"
★ 65Robot_Navigation_RL. Two paper About robot navigation in dynamic environment
★ 46CONVERT. [ACM MM 2023] An official source code for paper "CONVERT: Contrastive Graph Clustering with Reliable Augmentation".
★ 50CCGC. [AAAI 2023] An official source code for paper Cluster-guided Contrastive Graph Clustering Network.
★ 97ICCV23-IDPT. The code for the paper "Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud Models" (ICCV'23).
★ 111BadHash. The official implementation of BadHash
★ 58AdvCLIP. The implementation of our ACM MM 2023 paper "AdvCLIP: Downstream-agnostic Adversarial Examples in Multimodal Contrastive Learning"
★ 100Awesome-Nighttime-Enhancement. Collection of recent nighttime enhancement works, including papers, codes, datasets, and metrics.
★ 122FogRemoval. [ACCV22] Structure Representation Network and Uncertainty Feedback Learning for Dense Non-Uniform Fog Removal, https://arxiv.org/abs/2210.03061
★ 164UAL. The code for ECCV2022 paper: Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval
★ 55ISR_ICCV2023_Oral. The code for ICCV2023 Oral paper: Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-identification
★ 82MG-SCR. [IJCAI-2021] Official Codes for "Multi-Level Graph Encoding with Structural-Collaborative Relation Learning for Skeleton-Based Person Re-Identification"
★ 57SGE-LA. [IJCAI-2020] Official Codes for "Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification"
★ 69Locality-Awareness-SGE. [TPAMI-2022] Official Codes for "A Self-Supervised Gait Encoding Approach with Locality-Awareness for 3D Skeleton Based Person Re-Identification"
★ 78TranSG. [CVPR-2023] Official Codes for "TranSG: Transformer-Based Skeleton Graph Prototype Contrastive Learning with Structure-Trajectory Prompted Reconstruction for Person Re-Identification"
★ 98SM-SGE. [ACMMM-2021] Official Codes for "SM-SGE: A Self-Supervised Multi-Scale Skeleton Graph Encoding Framework for Person Re-Identification"
★ 57SimMC. [IJCAI-2022] Official Codes for "SimMC: Simple Masked Contrastive Learning of Skeleton Representations for Unsupervised Person Re-Identification"
★ 59Hi-MPC. [IJCV-2024] Official Codes for "Hierarchical Skeleton Meta-Prototype Contrastive Learning with Hard Skeleton Mining for Unsupervised Person Re-Identification"
★ 67NvEM. [ACM MM 2021 Oral] Official repo of "Neighbor-view Enhanced Model for Vision and Language Navigation"
★ 78ConsistentTeacher. [CVPR2023 Highlight] Consistent-Teacher: Towards Reducing Inconsistent Pseudo-targets in Semi-supervised Object Detection
★ 315SigmaCCS. [Communications Chemistry 2023] Highly accurate and large-scale collision cross section prediction with graph neural network for compound identification
★ 61Openvino_Yolov5_async. yolov5的openvino模型,带异步推理
★ 57ScaleVLN. [ICCV 2023 Oral]: Scaling Data Generation in Vision-and-Language Navigation
★ 225DTP. [ICCV 2023] Official implement of <Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement>
★ 71RFN. [TCSVT2022] Industria Scene Text Detection
★ 84RelightHuman.
★ 55SiT. Official implementation of "Self-slimmed Vision Transformer" (ECCV2022)
★ 72nighttime_dehaze. [ACMMM2023] "Enhancing Visibility in Nighttime Haze Images Using Guided APSF and Gradient Adaptive Convolution", https://arxiv.org/abs/2308.01738
★ 185S-Aware-network. [AAAI23] Estimating Reflectance Layer from A Single Image: Integrating Reflectance Guidance and Shadow/Specular Aware Learning, https://arxiv.org/abs/2211.14751
★ 89MomentDiff. MomentDiff: Generative Video Moment Retrieval from Random to Real--NeurIPS 2023
★ 80SVFormer. Python
★ 90SimDA. [CVPR 2024] SimDA: Simple Diffusion Adapter for Efficient Video Generation
★ 128Awesome-Video-Diffusion-Models. [CSUR] A Survey on Video Diffusion Models
★ 2.3kAdvEncoder. The implementation of our ICCV 2023 paper "Downstream-agnostic Adversarial Examples"
★ 69VTSNN. Public code for VTSNN: A Virtual Temporal Spiking Neural Network (Fron. Neur.)
★ 53ASDA. Python
★ 86BISKD. C++
★ 8CRAC. Python
★ 13ETRIS. [ICCV-2023] The official code of Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation
★ 138BISSG. code for paper IJCAI2022
★ 13NeMF. [ICCV 2023] NeMF: Inverse Volume Rendering with Neural Microflake Field
★ 68CSWinTT. Transformer Tracking with Cyclic Shifting Window Attention (CSWinTT)
★ 69MHFormer. [CVPR 2022] MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
★ 612C3BN. accepted by ICME2023 oral(CCF B)
★ 64DiffusionTrack. [AAAI 2024] DiffusionTrack: Diffusion Model For Multi-Object Tracking. DiffusionTrack is the first work to employ the diffusion model for multi-object tracking by formulating it as a generative noise-to-tracking diffusion process.
★ 206AAAI2023-PVD. [IJCV] Official Implementation of PVD and PVDAL: http://sk-fun.fun/PVD-AL/
★ 180PBRNet. The code is for PBRnet for action detection
★ 76FIMVC-VIA. [IEEE TNNLS 2022] Code of "Fast Incomplete Multi-view Clustering with View-independent Anchors"
★ 69LVOS.
★ 92EOMSC-CA. [AAAI 2022] Code of "Efficient One-pass Multi-view Subspace Clustering with Consensus Anchors"
★ 83CoDT. The implementaion of CoDT on the task of NTU-60+->PKUMMD
★ 76Variational-Degeneration-to-Structural-Refinement.
★ 54UniPC. [NeurIPS 2023] UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Models
★ 357Math-Reasoning-With-PLMs. EMNLP22: Multi-View Reasoning: Consistent Contrastive Learning for Math Word Problem
★ 218SnowFormer. This is the official implementation code repository of SnowFormer: Scale-aware Transformer via Context Interaction for Single Image Desnowing
★ 92UDR-S2Former_deraining. [ICCV'23] Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain Streaks
★ 149DC-ShadowNet-Hard-and-Soft-Shadow-Removal. [ICCV2021]"DC-ShadowNet: Single-Image Hard and Soft Shadow Removal Using Unsupervised Domain-Classifier Guided Network", https://arxiv.org/abs/2207.10434
★ 255DiffMIC. [MICCAI 2023] DiffMIC: Dual-Guidance Diffusion Network for Medical Image Classification
★ 173ViWS-Net. [ICCV 2023] Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial Backpropagation
★ 66night-enhancement. [ECCV2022] "Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression", https://arxiv.org/abs/2207.10564
★ 465SIGA. [CVPR2023] Self-supervised Implicit Glyph Attention for Text Recognition
★ 110RecommendFlow. A Industrialization recommendation system which can easy to build pipeline between local/hdfs train data and tensorflow2.x model.
★ 66VideoDesnowing. [ICCV 2023] Snow Removal in Video: A New Dataset and A Novel Method
★ 71FaissSearcher. a common faiss searcher
★ 169MaskedDenoising. [CVPR 2023] Masked Image Training for Generalizable Deep Image Denoising https://arxiv.org/abs/2303.13132
★ 285LoGoNet. [CVPR2023] LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross-Modal Fusion
★ 284HoP. [ICCV 2023] Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction
★ 196Co-DETR. [ICCV 2023] DETRs with Collaborative Hybrid Assignments Training
★ 1.4kMutexMatch4SSL. "MutexMatch: Semi-Supervised Learning with Mutex-Based Consistency Regularization" by Yue Duan (TNNLS)
★ 72GAC. Offical implementation of "Gated Attention Coding for Training High-performance and Efficient Spiking Neural Networks" (AAAI2024)
★ 121VQ-Font. [ICCV 2023] Few shot font generation via transferring similarity guided global and quantization local styles
★ 154ICLR_TINY_SNN. Offical implementation of "When Spiking Neural Networks Meet Temporal Attention Image Decoding And AdaptiveE Spiking Neuron" (ICLR2023)
★ 67VPD. [ICCV 2023] VPD is a framework that leverages the high-level and low-level knowledge of a pre-trained text-to-image diffusion model to downstream visual perception tasks.
★ 540Exp-BLIP. Official implementation of BMVC2023 Oral paper: 《Describe Your Facial Expressions by Linking Image Encoders and Large Language Models》
★ 65Parameterized-Cost-Volume-for-Stereo-Matching. PCVNet (ICCV2023)
★ 89Data-Copilot. Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
★ 1.5kNVDS. ICCV 2023 "Neural Video Depth Stabilizer" (NVDS) & TPAMI 2024 "NVDS+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation" (NVDS+)
★ 528SPC. [BMVC2023] Spatial and Planar Consistency for Semi-Supervised Volumetric Medical Image Segmentation
★ 77XNet. [ICCV2023] XNet: Wavelet-Based Low and High Frequency Merging Networks for Semi- and Supervised Semantic Segmentation of Biomedical Images
★ 215UniVTG. [ICCV 2023] UniVTG: Towards Unified Video-Language Temporal Grounding
★ 380DS2DP. The official implement of DS2DP [TGRS 2022]
★ 63On-the-fly-Category-Discovery. Code release for Your “On-the-fly Category Discovery (CVPR 2023)”
★ 58ProST. Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval --ICCV2023 Oral
★ 92Partial2Complete. [ICCV 2023] P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds
★ 194stock. stock股票.获取股票数据,计算股票指标,筹码分布,识别股票形态,综合选股,选股策略,股票验证回测,股票自动交易,支持PC及移动设备。
★ 14kETPNav. [TPAMI 2024] Official repo of "ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments"
★ 478BACL. Balanced Classification: A Unified Framework for Long-Tailed Object Detection (TMM 2023)
★ 101ViECap. Transferable Decoding with Visual Entities for Zero-Shot Image Captioning, ICCV 2023
★ 167PCS. Boosting Single Image Super-Resolution via Partial Channel Shifting.
★ 77DeRy. [NeurIPS2022] Deep Model Reassembly
★ 256DDS2M. The official implementation of DDS2M [ICCV 2023].
★ 126IAGNet. The Pytorch implementation of Grounding 3D Object Affordance from 2D Interactios in Images.
★ 138alice. Python
★ 58CCD. [ICCV2023] Self-supervised Character-to-Character Distillation for Text Recognition
★ 1533D_Semantic_Subspace_Traverser. C++
★ 64DT-ST. Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic Segmentation(CVPR-2023)
★ 81CASE. Accepted by ICCV2023, Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based Approach
★ 106PRG4SSL-MNAR. "Towards Semi-supervised Learning with Non-random Missing Labels" by Yue Duan (ICCV 2023)
★ 76zj-admin. vue3 ts vite elementplus nestjs mongoose zhangjincli docker
★ 104maozi-cloud-parent. 【脚手架】基于 SpringCloud Alibaba Dubbo 二开封装
★ 1.4kCode-Nest. springboot 项目 Code-Nest 程序员一体化社区 学习项目
★ 773pptshow. Java generates PPT documents and supports the new features of PPTX version 2010 / Java生成PPT文档,支持2010版PPTX新特性
★ 398BiliBili-Lucky-Draw. B站抽奖转发——薅羊毛脚本 : 一个小脚本能够帮助你去看看B站上面今天有哪些Up有抽奖活动,然后还能帮助你自动进行抽奖(转发动态+关注),毕竟抽奖总得试试吗,万一中奖了呢
★ 536phpzlc. PHPZlc是基于Symfony的脚手架工具
★ 216multiPrime. multiPrime is a mismatch-tolerant minimal primer set design tool for large and diverse sequences (e.g. Virus). Here is a web-based version (test: http://multiPrime.cn)
★ 456CrossThreadQueue. A thread safe queue that can be used in multi-thread project
★ 26geaflow. Apache GeaFlow: A Streaming Graph Computing Engine.
★ 788nv21FastConverter4RK. nv21 converter very fast version
★ 31qqrobot-sdk. QQ机器人一站式开发框架
★ 249VLN-BEVBert. [ICCV 2023] Official repo of "BEVBert: Multimodal Map Pre-training for Language-guided Navigation"
★ 260TF-ICON. [ICCV 2023] "TF-ICON: Diffusion-Based Training-Free Cross-Domain Image Composition" (Official Implementation)
★ 814VCSR. [ACM MM 2023] Visual Causal Scene Refinement for Video Question Answering
★ 2SkeletonMAE. SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training
★ 86CausalVLR. Visual-Linguistic Causal Learning Open-source Framework
★ 2CMCIR. [IEEE T-PAMI 2023] Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering
★ 78CMCIR. [IEEE T-PAMI 2023] Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering
★ 20TSTDDs. [IEEE SPL 2018] Global Temporal Representation based CNNs for Infrared Action Recognition
★ 7MIR. [MIR 2022] Causal Reasoning Meets Visual Representation Learning: A Prospective Study
★ 3CDFAG. Transferable Feature Representation for Visible-to-Infrared Cross-Dataset Human Action Recognition (Complexity 2018)
★ 13VLCI. Implicit Visual-Linguistic Deconfounding for Radiology Report Generation
★ 2HORLN. Python
★ 14CMCRL. The official implementation of “Cross-Modal Causal Representation Learning for Radiology Report Generation” (IEEE T-IP 2025)
★ 68JSRDA. [IEEE T-CSVT 2019] Hierarchically Learned View-Invariant Representations for Cross-View Action Recognition
★ 14DIVAFN. [IEEE T-IP 2020] Deep Image-to-Video Adaptation and Fusion Networks for Action Recognition
★ 7SAKDN. [IEEE T-IP 2021] Semantics-aware Adaptive Knowledge Distillation for Cross-modal Action Recognition
★ 29TCGL. [IEEE T-IP 2022] TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning
★ 24HORLN. [IEEE T-II 2022] Hybrid-Order Representation Learning for Electricity Theft Detection
★ 1SAKDN. [IEEE T-IP 2021] Semantics-aware Adaptive Knowledge Distillation for Cross-modal Action Recognition
★ 1TCGL. [IEEE T-IP 2022] TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning
★ 1gpt_academic. 为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, moss等。
★ 71kHORLN. Hybrid-Order Representation Learning for Electricity Theft Detection (IEEE T-II 2022)
★ 2GATNE. Source code and dataset for KDD 2019 paper "Representation Learning for Attributed Multiplex Heterogeneous Network"
★ 536awesome-action-recognition. A curated list of action recognition and related area resources
★ 4kadda_mnist64. Adversarial Discriminative Domain Adaptation with MNIST 64x64 in Lasagne-Theano
★ 33