This is your work, valued
Awesome-Colorful-LLM. Recent advancements propelled by large language models (LLMs), encompassing an array of domains including Vision, Audio, Agent, Robotics, Fundamental Sciences such as Mathematics, and Ominous.
128Streaming-Grounded-SAM-2. Grounded Tracking for Streaming Videos
127Awesome-Multimodal-Memory. [TMLR 2025] Reading List of Memory Augmented Multimodal Research, including multimodal context modeling, memory in vision and robotics, and external memory/knowledge augmented MLLM.
70VideoHallucer. VideoHallucer, The first comprehensive benchmark for hallucination detection in large video-language models (LVLMs)
43LM-Research-Hub. Language Modeling Research Hub, a comprehensive compendium for enthusiasts and scholars delving into the fascinating realm of language models (LMs), with a particular focus on large language models (LLMs)
19VSTAR. [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information
16NLPCC-2022-Shared-Task-4. Multimodal Dialogue Understanding and Generation
6CDBert. [ACL2023] Shuo Wen Jie Zi is a new learning paradigm that enhances the semantics understanding ability of the Chinese PLMs with dictionary knowledge and structure of Chinese characters
5MM-NIAVH. Pressure Testing Large Video-Language Models (LVLM): Doing multimodal retrieval from LVLM at any video lengths to measure accuracy
4Hallucination. A reading list of hallucination in Generative Models
1MARL_SG. [EMNLP2022] We propose a new collaborative reasoning method on mutli-modal graphs for multimodal dialogue
1FlowExtractor. https://github.com/v-iashin/video_features --> add script for flow extraction
1