This is your work, valued
Ph.D. student @ Westlake University & Zhejiang University, and research in the field of AI.
DyCoke. [CVPR 2025] DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
114OmniZip. [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
102LVOmniBench. LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
41VidKV. VidKV: Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
25OmniAgent. OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding
22MGFR. [ICLR 2025 Spotlight] Overcoming False Illusions in Real-World Face Restoration with Multi-Modal Guided Diffusion Model
16KD-TAO.github.io. Academic Personal Homepage of Keda Tao
2