founding rs. cs phd.
UnIVAL. [TMLR23] Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.
236ViCHA. [BMVC22] Official Implementation of ViCHA: "Efficient Vision-Language Pretraining with Visual Concepts and Hierarchical Alignment"
54xl-vlms. XL-VLMs: General Repository for eXplainable Large Vision Language Models
52TFood. [CVPRW22] Official Implementation of T-Food: "Transformer Decoders with MultiModal Regularization for Cross-Modal Food Retrieval". Accepted at CVPR22 's MULA Workshop.
34eP-ALM. [ICCV23] Official implementation of eP-ALM: Efficient Perceptual Augmentation of Language Models.
27ima-lmms. [NeurIPS2024] Official code for (IMA) Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
23EvALign-ICL. [ICLR2024] (EvALign-ICL Benchmark) Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
22VLPCook. Official implementation of VLPCook: Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
16lerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
2