Research in YouTu Tencent. swordli@tencent.com
Efficient-Multimodal-LLMs-Survey. Efficient Multimodal Large Language Models: A Survey
386Evaluation-Multimodal-LLMs-Survey. A Survey on Benchmarks of Multimodal Large Language Models
157lightDSFD. light DSFD
94Vote3Deep_lidar. Implementation of Vote3Deep algorithm on KITTI Object detection data to reproduce the current benchmark results.
43LLaVA-MR. LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
9Awesome-AI-Search. This repository offers an in-depth review of these advancements, focusing on Text-based AI Search, Web Browsing Agents, Multimodal AI Search, Benchmarks, Software, and Products.
4LLaVA-VSD. LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
2vgg11_SSD. Python
1cv2. 一杯淡茶,一首音乐,一句代码,一行文献,纵浪大化,不喜不惧
1ilsvrc_data. This code is used for ilsvrc data tar and lmdb
1