This is your work, valued
ProST. Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval --ICCV2023 Oral
★ 92MomentDiff. MomentDiff: Generative Video Moment Retrieval from Random to Real--NeurIPS 2023
★ 80NASA. Neighborhood-Adaptive Structure Augmented Metric Learning -- AAAI2022 oral
★ 8DKPH. Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval ECCV22
★ 7DFRQ. Deep Fourier Ranking Quantization for Semi-Supervised Image Retrieval -- TIP22
★ 6Awesome-Incremental-Learning. Awesome Incremental Learning
★ 4.5kebooks.
★ 158vstar. PyTorch Implementation of "V* : Guided Visual Search as a Core Mechanism in Multimodal LLMs"
★ 708Awesome_Long_Form_Video_Understanding. Awesome papers & datasets specifically focused on long-term videos.
★ 381VTimeLLM. [CVPR'2024 Highlight] Official PyTorch implementation of the paper "VTimeLLM: Empower LLM to Grasp Video Moments".
★ 295MESM. The official code of Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval (AAAI2024)
★ 32Bunny. A family of lightweight multimodal models.
★ 1.1kAwesome-LLMs-for-Video-Understanding. 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
★ 3.3kvicreg. VICReg official code base
★ 577PaperReading. Paper Reading of IMCC groups.
★ 18Awesome-Video-Diffusion. A curated list of recent diffusion models for video generation, editing, and various other applications.
★ 5.7kCTP. [ICCV2023] - CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation
★ 38VALOR. [TPAMI2024] Codes and Models for VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
★ 311bolei_awesome_posters. CVPR and NeurIPS poster examples and templates
★ 2kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kVAST. [NIPS2023] Code and Model for VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
★ 302ECLIPSE. Python
★ 33DeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43klearning_research. 本人的科研经验
★ 14kVisCPM. [ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列
★ 1.1kDFRQ. Deep Fourier Ranking Quantization for Semi-Supervised Image Retrieval -- TIP22
★ 6ONE-PEACE. A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
★ 1.1kMomentDiff. MomentDiff: Generative Video Moment Retrieval from Random to Real--NeurIPS 2023
★ 80SigmaCCS. [Communications Chemistry 2023] Highly accurate and large-scale collision cross section prediction with graph neural network for compound identification
★ 61Openvino_Yolov5_async. yolov5的openvino模型,带异步推理
★ 57DTP. [ICCV 2023] Official implement of <Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement>
★ 71RFN. [TCSVT2022] Industria Scene Text Detection
★ 84SiT. Official implementation of "Self-slimmed Vision Transformer" (ECCV2022)
★ 72SVFormer. Python
★ 90SimDA. [CVPR 2024] SimDA: Simple Diffusion Adapter for Efficient Video Generation
★ 128Awesome-Video-Diffusion-Models. [CSUR] A Survey on Video Diffusion Models
★ 2.3kAdvEncoder. The implementation of our ICCV 2023 paper "Downstream-agnostic Adversarial Examples"
★ 69VTSNN. Public code for VTSNN: A Virtual Temporal Spiking Neural Network (Fron. Neur.)
★ 53ASDA. Python
★ 86SkeletonMAE. SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training
★ 86RelightHuman.
★ 55GraphMLP. [PR 2025] GraphMLP: A Graph MLP-Like Architecture for 3D Human Pose Estimation
★ 75NeMF. [ICCV 2023] NeMF: Inverse Volume Rendering with Neural Microflake Field
★ 68StridedTransformer-Pose3D. [TMM 2023] Exploiting Temporal Contexts with Strided Transformer for 3D Human Pose Estimation
★ 355CSWinTT. Transformer Tracking with Cyclic Shifting Window Attention (CSWinTT)
★ 69MHFormer. [CVPR 2022] MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
★ 612C3BN. accepted by ICME2023 oral(CCF B)
★ 64DiffusionTrack. [AAAI 2024] DiffusionTrack: Diffusion Model For Multi-Object Tracking. DiffusionTrack is the first work to employ the diffusion model for multi-object tracking by formulating it as a generative noise-to-tracking diffusion process.
★ 206AAAI2023-PVD. [IJCV] Official Implementation of PVD and PVDAL: http://sk-fun.fun/PVD-AL/
★ 180PBRNet. The code is for PBRnet for action detection
★ 76FIMVC-VIA. [IEEE TNNLS 2022] Code of "Fast Incomplete Multi-view Clustering with View-independent Anchors"
★ 69EOMSC-CA. [AAAI 2022] Code of "Efficient One-pass Multi-view Subspace Clustering with Consensus Anchors"
★ 83LVOS.
★ 92CoDT. The implementaion of CoDT on the task of NTU-60+->PKUMMD
★ 76Variational-Degeneration-to-Structural-Refinement.
★ 54UniPC. [NeurIPS 2023] UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Models
★ 357Math-Reasoning-With-PLMs. EMNLP22: Multi-View Reasoning: Consistent Contrastive Learning for Math Word Problem
★ 218SnowFormer. This is the official implementation code repository of SnowFormer: Scale-aware Transformer via Context Interaction for Single Image Desnowing
★ 92UDR-S2Former_deraining. [ICCV'23] Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain Streaks
★ 149DC-ShadowNet-Hard-and-Soft-Shadow-Removal. [ICCV2021]"DC-ShadowNet: Single-Image Hard and Soft Shadow Removal Using Unsupervised Domain-Classifier Guided Network", https://arxiv.org/abs/2207.10434
★ 255DiffMIC. [MICCAI 2023] DiffMIC: Dual-Guidance Diffusion Network for Medical Image Classification
★ 173ViWS-Net. [ICCV 2023] Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial Backpropagation
★ 66night-enhancement. [ECCV2022] "Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression", https://arxiv.org/abs/2207.10564
★ 465SIGA. [CVPR2023] Self-supervised Implicit Glyph Attention for Text Recognition
★ 110RecommendFlow. A Industrialization recommendation system which can easy to build pipeline between local/hdfs train data and tensorflow2.x model.
★ 66FaissSearcher. a common faiss searcher
★ 169VideoDesnowing. [ICCV 2023] Snow Removal in Video: A New Dataset and A Novel Method
★ 71MaskedDenoising. [CVPR 2023] Masked Image Training for Generalizable Deep Image Denoising https://arxiv.org/abs/2303.13132
★ 285LoGoNet. [CVPR2023] LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross-Modal Fusion
★ 284HoP. [ICCV 2023] Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction
★ 196Co-DETR. [ICCV 2023] DETRs with Collaborative Hybrid Assignments Training
★ 1.4kMutexMatch4SSL. "MutexMatch: Semi-Supervised Learning with Mutex-Based Consistency Regularization" by Yue Duan (TNNLS)
★ 72ProST. Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval --ICCV2023 Oral
★ 92Exp-BLIP. Official implementation of BMVC2023 Oral paper: 《Describe Your Facial Expressions by Linking Image Encoders and Large Language Models》
★ 65VPD. [ICCV 2023] VPD is a framework that leverages the high-level and low-level knowledge of a pre-trained text-to-image diffusion model to downstream visual perception tasks.
★ 540VQ-Font. [ICCV 2023] Few shot font generation via transferring similarity guided global and quantization local styles
★ 154ICLR_TINY_SNN. Offical implementation of "When Spiking Neural Networks Meet Temporal Attention Image Decoding And AdaptiveE Spiking Neuron" (ICLR2023)
★ 67GAC. Offical implementation of "Gated Attention Coding for Training High-performance and Efficient Spiking Neural Networks" (AAAI2024)
★ 121SPC. [BMVC2023] Spatial and Planar Consistency for Semi-Supervised Volumetric Medical Image Segmentation
★ 77Parameterized-Cost-Volume-for-Stereo-Matching. PCVNet (ICCV2023)
★ 89Data-Copilot. Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
★ 1.5kNVDS. ICCV 2023 "Neural Video Depth Stabilizer" (NVDS) & TPAMI 2024 "NVDS+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation" (NVDS+)
★ 528XNet. [ICCV2023] XNet: Wavelet-Based Low and High Frequency Merging Networks for Semi- and Supervised Semantic Segmentation of Biomedical Images
★ 215DS2DP. The official implement of DS2DP [TGRS 2022]
★ 63On-the-fly-Category-Discovery. Code release for Your “On-the-fly Category Discovery (CVPR 2023)”
★ 58Partial2Complete. [ICCV 2023] P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds
★ 194ETPNav. [TPAMI 2024] Official repo of "ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments"
★ 4783D_Semantic_Subspace_Traverser. C++
★ 64DT-ST. Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic Segmentation(CVPR-2023)
★ 81CCD. [ICCV2023] Self-supervised Character-to-Character Distillation for Text Recognition
★ 153PRG4SSL-MNAR. "Towards Semi-supervised Learning with Non-random Missing Labels" by Yue Duan (ICCV 2023)
★ 76alice. Python
★ 58IAGNet. The Pytorch implementation of Grounding 3D Object Affordance from 2D Interactios in Images.
★ 138CASE. Accepted by ICCV2023, Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based Approach
★ 106DDS2M. The official implementation of DDS2M [ICCV 2023].
★ 126DeRy. [NeurIPS2022] Deep Model Reassembly
★ 256PCS. Boosting Single Image Super-Resolution via Partial Channel Shifting.
★ 77ViECap. Transferable Decoding with Visual Entities for Zero-Shot Image Captioning, ICCV 2023
★ 167CausalVLR. CausalVLR: A Toolbox and Benchmark for Vision-Language Causal Reasoning (多模态因果推理开源框架)
★ 1.1kTF-ICON. [ICCV 2023] "TF-ICON: Diffusion-Based Training-Free Cross-Domain Image Composition" (Official Implementation)
★ 814VLN-BEVBert. [ICCV 2023] Official repo of "BEVBert: Multimodal Map Pre-training for Language-guided Navigation"
★ 260DiffusionRet. [ICCV 2023] DiffusionRet: Generative Text-Video Retrieval with Diffusion Model
★ 141UniVTG. [ICCV 2023] UniVTG: Towards Unified Video-Language Temporal Grounding
★ 380DORi. Public repository for DORi: Discovering Object Relationships for Moment Localization of a Natural Language Query in a Video Code accompanying the paper
★ 21LGI4temporalgrounding. Repository for the CVPR-20 paper "Local-Global Video-Text Interactions for Temporal Grounding"
★ 132MS-DETR. An official implementation for MS-DETR in ACL'23
★ 17ShufflingVideosForTSG. Code for ECCV 2022 paper "Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding"
★ 29ustcbeamer. USTC Beamer 模板(基于学校公用 PPT 模板)
★ 347opencompass. OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
★ 7.3kChatGLM2-6B. ChatGLM2-6B: An Open Bilingual Chat LLM | 开源双语对话语言模型
★ 16kVideo-LLaMA. [EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
★ 3.1kLLaMA-Adapter. [ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
★ 5.9kvisprog. Official code for VisProg (CVPR 2023 Best Paper!)
★ 774DeepLearning-500-questions. 深度学习500问,以问答形式对常用的概率知识、线性代数、机器学习、深度学习、计算机视觉等热点问题进行阐述,以帮助自己及有需要的读者。 全书分为18个章节,50余万字。由于水平有限,书中不妥之处恳请广大读者批评指正。 未完待续............ 如有意合作,联系scutjy2015@163.com 版权所有,违权必究 Tan 2018.06
★ 58kCausal_Video_Moment_Retrieval. The codes and features of the re-implementation of SIGIR 2021 work "Deconfounded Video Moment Retrieval with Causal Intervention"
★ 36LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kInternVideo. [ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2.3kmPLUG-Owl. mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
★ 2.5kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kMacaw-LLM. Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration
★ 1.6kBACL. Balanced Classification: A Unified Framework for Long-Tailed Object Detection (TMM 2023)
★ 101SparseR-CNN. [CVPR2021, PAMI2023] End-to-End Object Detection with Learnable Proposal
★ 1.3kNASA. Neighborhood-Adaptive Structure Augmented Metric Learning -- AAAI2022 oral
★ 8Awesome-Trajectory-Motion-Prediction-Papers.
★ 1.1kCompositional-Temporal-Grounding.
★ 31vqvae. A pytorch implementation of the vector quantized variational autoencoder (https://arxiv.org/abs/1711.00937)
★ 901pytorch-vqvae. Vector Quantized VAEs - PyTorch Implementation
★ 955SeqTR. SeqTR: A Simple yet Universal Network for Visual Grounding
★ 144unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22k