This is your work, valued
PhD Candidate of University of Science and Technology of China. I’m currently working on Multimodal AI.
TextCoT. [ACM TOMM] Official implementation of "TextCoT: Zoom-In for Enhanced Multimodal Text-Rich Image Understanding"
AdaptPrune. The official repo for “Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models”.