Ph.D. student. Research Interests: LLM-Agents, Vision-Language.
Awesome-Embodied-Robotics-and-Agent. This is a curated list of "Embodied AI or robot with Large Language Models" research. Watch this repository for the latest updates! 🔥
1.8kS2-Transformer. [IJCAI 2022] Official Pytorch code for paper “S2 Transformer for Image Captioning”
86PKOL. [TIP 2022] Official code of paper “Video Question Answering with Prior Knowledge and Object-sensitive Learning”
46Multi-Modal-Large-Language-Learning. Awesome multi-modal large language paper/project, collections of popular training strategies, e.g., PEFT, LoRA.
27GLSCL. [TIP25] Code for "Text-Video Retrieval with Global-Local Semantic Consistent Learning"
16VCRN. Python
11SPT. [TCSVT23] Official code for "SPT: Spatial Pyramid Transformer for Image Captioning".
10Awesome-Post-Training-for-MLLMs. A curated list of "A Survey on Post-training of Multimodal Large Language Models" research. Watch this repository for latest updates! 🔥
10OmniCharacter-plus. [TPAMI26] Official codebase for "OmniCharacter++: Towards Comprehensive Benchmark for Realistic Role-Playing Agents" 🔥
93D-Vision-and-Language. Collection of recent 3D Vision and Language research
8SNLC. [PR23] The implementation of the paper ''Learning Visual Question Answering on Controlled Semantic Noisy Labels''
8OmniCharacter. [ACL25] Official codebase for "OmniCharacter:Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction" 🔥
8UMP_TVR. [TCSVT24] The implementation of paper "UMP: Unified Modality-aware Prompt Tuning for Text-Video Retrieval".
7zchoi.
7DAST. [MM23] Code for paper "Depth-Aware Sparse Transformer for Video-Language Learning"
6MAN. Python
6videoqa_model. Jupyter Notebook
5RSTNet. RSTNet: Captioning with Adaptive Attention on Visual and Non-Visual Words (CVPR 2021)
5VQAC. Python
5PRL. The code for the ICLR25 submission paper
3sam. SAM: Sharpness-Aware Minimization (PyTorch)
2Vision-and-Language-Benchmark. Codebase for research of vision&language, including various multimodal task pipline (e.g., image captioning, VQA, video-text retrieval), customizable dataset (e.g., MS-COCO, ActivityNet, MSR-VTT), pre-trained model acquire (e.g., CLIP, BLIP-2)
2McQuic. Repository of CVPR'22 paper "Unified Multivariate Gaussian Mixture for Efficient Neural Image Compression"
2EMCL. [NeurIPS 2022] Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations
2metrics. 📊 An infographics generator with 30+ plugins and 200+ options to display stats about your GitHub account and render them as SVG, Markdown, PDF or JSON!
2rich. Rich is a Python library for rich text and beautiful formatting in the terminal.
2LMaaS-Papers. Awesome papers on Language-Model-as-a-Service (LMaaS)
2Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
1Awesome-Weak-to-Strong-Generalization. 🔥🔥🔥 This repository curates research on Weak-to-Strong Generalization across LLMs, multimodal learning, and beyond, focusing on how strong models learn from weak supervision and surpass their teachers. Stay tuned for the latest updates!
1