This is your work, valued
mllms_know. [ICLR'25] Official code for the paper 'MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs'
382text-based-traffic-understanding. The code and dataset of the KDD23 paper 'A Study of Situational Reasoning for Traffic Understanding'
18visual_crop_zsvqa. Jupyter Notebook
12mllm-perceptual-limitation. Code and data for paper 'Exploring Perceptual Limitation of Multimodal Large Language Models'
11commonsense-with-KG. Code for the paper "An Empirical Investigation of Commonsense Self-Supervision with Knowledge Graphs"
4zjrr.github.io. HTML
2Awesome_Think_With_Images. Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
1PerLim. The code and dataset of the PerLim paper
1old. Github Pages template for academic personal websites, forked from mmistakes/minimal-mistakes
1saccharomycetes.github.io. HTML
1text-based-traffic.
1Continuous_Thinking_Modeling. Code space for continuous thinking modeling project
1