This is your work, valued
EditThinker. Unlocking Iterative Reasoning for Any Image Editor
112AL-Ref-SAM2. [AAAI 2025] AL-Ref-SAM 2: Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
93LLaVA-ST. [CVPR 2025] LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
84Temporal-R1. Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
62