This is your work, valued
UniversalRAG. [ACL 2026 Oral] UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
174Video-Oasis. [ECCV 2026] Video-Oasis: Rethinking Evaluation of Video Understanding
33MAGIC-video. Python
6SceneGraphVLM. Jupyter Notebook
11WorldMM. [CVPR 2026 Highlight] WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
99ProVideLLM. [ICCV 2025] Streaming VideoLLMs for Real-time Procedural Video Understanding
20DecAF. [ICLR 2026] Official implementation of "Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation"
36STTM. [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
61Seen_to_Scene. [CVPR 2026 Findings Paper] Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting
7CMTM. [ICIP 2025 Oral Paper] CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
7StreamChat. Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025
111PAWS. Official code for "Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning"
3OTT-Vid. Official code for "OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models"
9Awesome-Streaming-Video-Understanding. 🔥🔥🔥 [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding — Building Always-on, Real-time Video AI 🤖
421aurora. Implementation of the Aurora model for Earth system forecasting
978CDPruner. [NeurIPS 2025] Official code for paper: Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.
106Tango. Repo for paper "Tango: Taming Visual Signals for Efficient Video Large Language Models"
9V-CAST. V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
34FrameFusion. [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"
75AgilePruner. [ICLR 2026] AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
28Awesome-Multimodal-Token-Compression. [TMLR 2026] Survey: https://arxiv.org/pdf/2507.20198
376TimeLens. [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
163Nuwa. Official Reop of Nüwa: Mending the Spatial Integrity Torn by LVLM Acceleration [ICLR26]
6OTPrune. Official code for the paper: OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
6AOT. [CVPR 2026 🎉] Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
15Awesome-Multimodal-Object-Tracking. A continuously updated project to track the latest progress in the field of multi-modal object tracking. This project focuses solely on single-object tracking.
1.1kFlashVID. [ICLR 2026 Oral] FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
116GenCLIP. GenCLIP (Pattern Recognition, Volume 178, October 2026, 113406)
3CPLVAD. This repository contains the implementation of CPLVAD (2026 ICASSP).
7Awesome-Token-Compress. A paper list of some recent works about Token Compress for Vit and VLM
944AKS. [CVPR 2025] Adaptive Keyframe Sampling for Long Video Understanding
228DisTime. DisTime: Distribution-based Time Representation for Video Large Language Models.
21opencode. The open source coding agent.
192kRT-DETRv4. [ECCV 2026] RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
564Awesome-Efficient-Arch. Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
407SwiftVGGT. [CVPR 2026 Findings] SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
95sam3. The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
11kGroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
10kAwesome-Scene-Graph-Generation. This is a repository for listing papers on scene graph generation and application.
706vjepa2. PyTorch code and models for VJEPA2 self-supervised learning from video.
4.4kQwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
4.1kDeepSeek-V3. Python
104knotion-mcp-server. Official Notion MCP Server
4.6kUnAV. Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline (CVPR 2023)
73AVicuna. [AAAI 2025] Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
34perception_test. Jupyter Notebook
255jepa. PyTorch code and models for V-JEPA self-supervised learning from video.
4.1kImageBind. ImageBind One Embedding Space to Bind Them All
9.1kAwesome-LLMs-for-Video-Understanding. 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
3.3kDiGIT. [CVPR 2025] Official implementation of the paper "DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer"
32STOV-TAL. [WACV-2025] Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
17Awesome-Temporal-Action-Detection-Temporal-Action-Proposal-Generation. Temporal Action Detection & Weakly Supervised Temporal Action Detection & Temporal Action Proposal Generation
591CoMoGaussian. [ICCV 2025] CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images
57UVCOM. [CVPR 2024] Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection
118DAMSDet. Python
73ProST. Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval --ICCV2023 Oral
92Keyword-DETR. Official Repository for "Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection" (AAAI 2025)
15VideoLLaMA3. Frontier Multimodal Foundation Models for Image and Video Understanding
1.2kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
18kTarDAL. CVPR 2022 | Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection.
206DsHmp. [CVPR-2024] Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation
83ViT-Adapter. [ICLR 2023 Spotlight] Vision Transformer Adapter for Dense Predictions
1.5kDeepSeek-R1.
92kDeformable-DETR. Deformable DETR: Deformable Transformers for End-to-End Object Detection.
4kPNAS-MOT. RAL 2024: PNAS-MOT: Multi-Modal Object Tracking with Pareto Neural Architecture Search
13VirConv. Virtual Sparse Convolution for Multimodal 3D Object Detection
384DMFormer. [IEEE T-CSVT] Decoupled Multimodal Transformers (DMFormer) for Referring Video Object Segmentation
5TaskWeave. [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
30TR-DETR. Official pytorch repository for "TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection" (AAAI 2024 Paper)
57QD-DETR. Official pytorch repository for "QD-DETR : Query-Dependent Video Representation for Moment Retrieval and Highlight Detection" (CVPR 2023 Paper)
251DeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
43kflash-attention. Fast and memory-efficient exact attention
25kmr-Blip. Official Implementation of "Chrono: A Simple Blueprint for Representing Time in MLLMs"
95InternVideo. [ECCV2024] Video Foundation Models & Data for Multimodal Understanding
2.3kByteTrack. [ECCV 2022] ByteTrack: Multi-Object Tracking by Associating Every Detection Box
6.6ktc-clip. [ECCV 2024] Official PyTorch implementation of TC-CLIP "Leveraging Temporal Contextualization for Video Action Recognition"
102MCTrack. [IROS2025]This is the offical implementation of the paper "MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving"
256Video-Swin-Transformer. This is an official implementation for "Video Swin Transformers".
1.7kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
11kCLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
34k3D-Detection-Tracking-Viewer. 3D detection and tracking viewer (visualization) for kitti & waymo dataset
5333D-Multi-Object-Tracker. A project for 3D multi-object tracking
383awesome-multiple-object-tracking. Resources for Multiple Object Tracking (MOT)
1.5kFairMOT. [IJCV-2021] FairMOT: On the Fairness of Detection and Re-Identification in Multi-Object Tracking
4.2kCGSTVG. [CVPR 2024] Context-Guided Spatio-Temporal Video Grounding
66MV-Adapter. An official pytorch implementation of the paper: [MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval].
14R2-Tuning. 🌀 R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding (ECCV 2024)
92groundingLMM. [CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.
965HERO_Video_Feature_Extractor. Video Feature Extraction Code for EMNLP 2020 paper "HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training"
118GSANet. [CVPR 2024] Guided Slot Attention for Unsupervised Video Object Segmentation
66mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
8.4kSMURF. [CVPRW 2025] SMURF: Continuous Dynamics for Motion-Deblurring Radiance Fields
25ICCV-2023-25-Papers. ICCV 2023-2025 Papers: Discover cutting-edge research from ICCV 2023-25, the leading computer vision conference. Stay updated on the latest in computer vision and deep learning, with code included. ⭐ support visual intelligence development!
968CVPR24-FAS. Official implementation of CVPR24 paper "Gradient Alignment for Cross-Domain Face Anti-Spoofing"
84Neural-Network-Diffusion. We introduce a novel approach for parameter generation, named neural network parameter diffusion (p-diff), which employs a standard latent diffusion model to synthesize a new set of parameters
886Depth-Anything. [CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation
8.2ks4. Structured state space sequence models
2.9kNATTEN. Fast Multi-dimensional Sparse Attention
780SegLossOdyssey. A collection of loss functions for medical image segmentation
4kScenePriors. Implementation of CVPR'23: Learning 3D Scene Priors with 2D Supervision
77External-Attention-pytorch. 🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐
12kDataloader-Optimization. Getting GPU Util 99%
33gigagan-pytorch. Implementation of GigaGAN, new SOTA GAN out of Adobe. Culmination of nearly a decade of research into GANs
1.9kFAPM_official. This repository contains the implementation of FAPM (2023 ICASSP).
25awesome-3d-reconstruction-papers. A collection of 3D reconstruction papers in the deep learning era.
910