This is your work, valued
VisionSelector. VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
FMAE. Python