This is your work, valued

Seoul

MINSEOK KANG

Advanced
@minseokii

UniversalRAG. [ACL 2026 Oral] UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities

174

Video-Oasis. [ECCV 2026] Video-Oasis: Rethinking Evaluation of Video Understanding

33

MAGIC-video. Python

6

SceneGraphVLM. Jupyter Notebook

11

WorldMM. [CVPR 2026 Highlight] WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning

99

ProVideLLM. [ICCV 2025] Streaming VideoLLMs for Real-time Procedural Video Understanding

20

DecAF. [ICLR 2026] Official implementation of "Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation"

36

STTM. [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

61

Seen_to_Scene. [CVPR 2026 Findings Paper] Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting

7

CMTM. [ICIP 2025 Oral Paper] CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation

7

StreamChat. Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025

111

PAWS. Official code for "Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning"

3

OTT-Vid. Official code for "OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models"

9

Awesome-Streaming-Video-Understanding. 🔥🔥🔥 [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding — Building Always-on, Real-time Video AI 🤖

421

aurora. Implementation of the Aurora model for Earth system forecasting

978

CDPruner. [NeurIPS 2025] Official code for paper: Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

106

Tango. Repo for paper "Tango: Taming Visual Signals for Efficient Video Large Language Models"

9

V-CAST. V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models

34

FrameFusion. [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"

75

AgilePruner. [ICLR 2026] AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models

28

Awesome-Multimodal-Token-Compression. [TMLR 2026] Survey: https://arxiv.org/pdf/2507.20198

376

TimeLens. [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

163

Nuwa. Official Reop of Nüwa: Mending the Spatial Integrity Torn by LVLM Acceleration [ICLR26]

6

OTPrune. Official code for the paper: OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport

6

AOT. [CVPR 2026 🎉] Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

15

Awesome-Multimodal-Object-Tracking. A continuously updated project to track the latest progress in the field of multi-modal object tracking. This project focuses solely on single-object tracking.

1.1k

FlashVID. [ICLR 2026 Oral] FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

116

GenCLIP. GenCLIP (Pattern Recognition, Volume 178, October 2026, 113406)

3

CPLVAD. This repository contains the implementation of CPLVAD (2026 ICASSP).

7

Awesome-Token-Compress. A paper list of some recent works about Token Compress for Vit and VLM

944

AKS. [CVPR 2025] Adaptive Keyframe Sampling for Long Video Understanding

228

DisTime. DisTime: Distribution-based Time Representation for Video Large Language Models.

21

opencode. The open source coding agent.

192k

RT-DETRv4. [ECCV 2026] RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models

564

Awesome-Efficient-Arch. Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

407

SwiftVGGT. [CVPR 2026 Findings] SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes

95

sam3. The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.

11k

GroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"

10k

Awesome-Scene-Graph-Generation. This is a repository for listing papers on scene graph generation and application.

706

vjepa2. PyTorch code and models for VJEPA2 self-supervised learning from video.

4.4k

Qwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.

4.1k

DeepSeek-V3. Python

104k

notion-mcp-server. Official Notion MCP Server

4.6k

UnAV. Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline (CVPR 2023)

73

AVicuna. [AAAI 2025] Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding

34

perception_test. Jupyter Notebook

255

jepa. PyTorch code and models for V-JEPA self-supervised learning from video.

4.1k

ImageBind. ImageBind One Embedding Space to Bind Them All

9.1k

Awesome-LLMs-for-Video-Understanding. 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.

3.3k

DiGIT. [CVPR 2025] Official implementation of the paper "DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer"

32

STOV-TAL. [WACV-2025] Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization

17

Awesome-Temporal-Action-Detection-Temporal-Action-Proposal-Generation. Temporal Action Detection & Weakly Supervised Temporal Action Detection & Temporal Action Proposal Generation

591

CoMoGaussian. [ICCV 2025] CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images

57

UVCOM. [CVPR 2024] Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

118

DAMSDet. Python

73

ProST. Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval --ICCV2023 Oral

92

Keyword-DETR. Official Repository for "Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection" (AAAI 2025)

15

VideoLLaMA3. Frontier Multimodal Foundation Models for Image and Video Understanding

1.2k

Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models

18k

TarDAL. CVPR 2022 | Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection.

206

DsHmp. [CVPR-2024] Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation

83

ViT-Adapter. [ICLR 2023 Spotlight] Vision Transformer Adapter for Dense Predictions

1.5k

DeepSeek-R1.

92k

Deformable-DETR. Deformable DETR: Deformable Transformers for End-to-End Object Detection.

4k

PNAS-MOT. RAL 2024: PNAS-MOT: Multi-Modal Object Tracking with Pareto Neural Architecture Search

13

VirConv. Virtual Sparse Convolution for Multimodal 3D Object Detection

384

DMFormer. [IEEE T-CSVT] Decoupled Multimodal Transformers (DMFormer) for Referring Video Object Segmentation

5

TaskWeave. [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection

30

TR-DETR. Official pytorch repository for "TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection" (AAAI 2024 Paper)

57

QD-DETR. Official pytorch repository for "QD-DETR : Query-Dependent Video Representation for Moment Retrieval and Highlight Detection" (CVPR 2023 Paper)

251

DeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

43k

flash-attention. Fast and memory-efficient exact attention

25k

mr-Blip. Official Implementation of "Chrono: A Simple Blueprint for Representing Time in MLLMs"

95

InternVideo. [ECCV2024] Video Foundation Models & Data for Multimodal Understanding

2.3k

ByteTrack. [ECCV 2022] ByteTrack: Multi-Object Tracking by Associating Every Detection Box

6.6k

tc-clip. [ECCV 2024] Official PyTorch implementation of TC-CLIP "Leveraging Temporal Contextualization for Video Action Recognition"

102

MCTrack. [IROS2025]This is the offical implementation of the paper "MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving"

256

Video-Swin-Transformer. This is an official implementation for "Video Swin Transformers".

1.7k

LAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence

11k

CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image

34k

3D-Detection-Tracking-Viewer. 3D detection and tracking viewer (visualization) for kitti & waymo dataset

533

3D-Multi-Object-Tracker. A project for 3D multi-object tracking

383

awesome-multiple-object-tracking. Resources for Multiple Object Tracking (MOT)

1.5k

FairMOT. [IJCV-2021] FairMOT: On the Fairness of Detection and Re-Identification in Multi-Object Tracking

4.2k

CGSTVG. [CVPR 2024] Context-Guided Spatio-Temporal Video Grounding

66

MV-Adapter. An official pytorch implementation of the paper: [MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval].

14

R2-Tuning. 🌀 R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding (ECCV 2024)

92

groundingLMM. [CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.

965

HERO_Video_Feature_Extractor. Video Feature Extraction Code for EMNLP 2020 paper "HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training"

118

GSANet. [CVPR 2024] Guided Slot Attention for Unsupervised Video Object Segmentation

66

mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377

8.4k

SMURF. [CVPRW 2025] SMURF: Continuous Dynamics for Motion-Deblurring Radiance Fields

25

ICCV-2023-25-Papers. ICCV 2023-2025 Papers: Discover cutting-edge research from ICCV 2023-25, the leading computer vision conference. Stay updated on the latest in computer vision and deep learning, with code included. ⭐ support visual intelligence development!

968

CVPR24-FAS. Official implementation of CVPR24 paper "Gradient Alignment for Cross-Domain Face Anti-Spoofing"

84

Neural-Network-Diffusion. We introduce a novel approach for parameter generation, named neural network parameter diffusion (p-diff), which employs a standard latent diffusion model to synthesize a new set of parameters

886

Depth-Anything. [CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation

8.2k

s4. Structured state space sequence models

2.9k

NATTEN. Fast Multi-dimensional Sparse Attention

780

SegLossOdyssey. A collection of loss functions for medical image segmentation

4k

ScenePriors. Implementation of CVPR'23: Learning 3D Scene Priors with 2D Supervision

77

External-Attention-pytorch. 🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐

12k

Dataloader-Optimization. Getting GPU Util 99%

33

gigagan-pytorch. Implementation of GigaGAN, new SOTA GAN out of Adobe. Culmination of nearly a decade of research into GANs

1.9k

FAPM_official. This repository contains the implementation of FAPM (2023 ICASSP).

25

awesome-3d-reconstruction-papers. A collection of 3D reconstruction papers in the deep learning era.

910