Xingjian Diao is a Ph.D. student in Computer Science at Dartmouth Collegeš², working on multimodal learning and reasoning.
SoundMind. We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow audio language models (ALMs) with deep bimodal reasoning abilities.
1.1kNAACL_2025_TWM. We introduce temporal working memory (TWM), which aims to enhance the temporal modeling capabilities of Multimodal foundation models (MFMs). This plug-and-play module can be easily integrated into existing MFMs. With our TWM, nine state-of-the-art models exhibit significant performance improvements across QA, captioning, and retrieval tasks.
315Survey4MusicAVQA. Survey for MusicAVQA
5Amuse. Amuse-Emnlp
3FT2TF. FT2TF
2Active-Learning-Annotation-Tool. An interactive annotation software which uses human-in-the-loop machine learning concepts known as āActive Learningā to label time sequences, in order to save a great deal of labelling expenses and take upon the challenges faced during the annotation process
2189VU_Github_sample. 189 VideoUnderstanding Github Sample
1