This is your work, valued

Yuxuan Wang

Expert
@patrick-tssn

No pride and no prejudice

Awesome-Colorful-LLM. Recent advancements propelled by large language models (LLMs), encompassing an array of domains including Vision, Audio, Agent, Robotics, Fundamental Sciences such as Mathematics, and Ominous.

128

Streaming-Grounded-SAM-2. Grounded Tracking for Streaming Videos

127

Awesome-Multimodal-Memory. [TMLR 2025] Reading List of Memory Augmented Multimodal Research, including multimodal context modeling, memory in vision and robotics, and external memory/knowledge augmented MLLM.

70

VideoHallucer. VideoHallucer, The first comprehensive benchmark for hallucination detection in large video-language models (LVLMs)

43

LM-Research-Hub. Language Modeling Research Hub, a comprehensive compendium for enthusiasts and scholars delving into the fascinating realm of language models (LMs), with a particular focus on large language models (LLMs)

19

VSTAR. [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information

16

NLPCC-2022-Shared-Task-4. Multimodal Dialogue Understanding and Generation

6

CDBert. [ACL2023] Shuo Wen Jie Zi is a new learning paradigm that enhances the semantics understanding ability of the Chinese PLMs with dictionary knowledge and structure of Chinese characters

5

MM-NIAVH. Pressure Testing Large Video-Language Models (LVLM): Doing multimodal retrieval from LVLM at any video lengths to measure accuracy

4

Hallucination. A reading list of hallucination in Generative Models

1

MARL_SG. [EMNLP2022] We propose a new collaborative reasoning method on mutli-modal graphs for multimodal dialogue

1

FlowExtractor. https://github.com/v-iashin/video_features --> add script for flow extraction

1