Sichuan ⇌ Italy

Haonan Zhang

Advanced
@zchoi

Ph.D. student. Research Interests: LLM-Agents, Vision-Language.

Awesome-Embodied-Robotics-and-Agent. This is a curated list of "Embodied AI or robot with Large Language Models" research. Watch this repository for the latest updates! 🔥

1.8k

S2-Transformer. [IJCAI 2022] Official Pytorch code for paper “S2 Transformer for Image Captioning”

86

PKOL. [TIP 2022] Official code of paper “Video Question Answering with Prior Knowledge and Object-sensitive Learning”

46

Multi-Modal-Large-Language-Learning. Awesome multi-modal large language paper/project, collections of popular training strategies, e.g., PEFT, LoRA.

27

GLSCL. [TIP25] Code for "Text-Video Retrieval with Global-Local Semantic Consistent Learning"

16

VCRN. Python

11

SPT. [TCSVT23] Official code for "SPT: Spatial Pyramid Transformer for Image Captioning".

10

Awesome-Post-Training-for-MLLMs. A curated list of "A Survey on Post-training of Multimodal Large Language Models" research. Watch this repository for latest updates! 🔥

10

OmniCharacter-plus. [TPAMI26] Official codebase for "OmniCharacter++: Towards Comprehensive Benchmark for Realistic Role-Playing Agents" 🔥

9

3D-Vision-and-Language. Collection of recent 3D Vision and Language research

8

SNLC. [PR23] The implementation of the paper ''Learning Visual Question Answering on Controlled Semantic Noisy Labels''

8

OmniCharacter. [ACL25] Official codebase for "OmniCharacter:Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction" 🔥

8

UMP_TVR. [TCSVT24] The implementation of paper "UMP: Unified Modality-aware Prompt Tuning for Text-Video Retrieval".

7

zchoi.

7

DAST. [MM23] Code for paper "Depth-Aware Sparse Transformer for Video-Language Learning"

6

MAN. Python

6

videoqa_model. Jupyter Notebook

5

RSTNet. RSTNet: Captioning with Adaptive Attention on Visual and Non-Visual Words (CVPR 2021)

5

VQAC. Python

5

PRL. The code for the ICLR25 submission paper

3

sam. SAM: Sharpness-Aware Minimization (PyTorch)

2

Vision-and-Language-Benchmark. Codebase for research of vision&language, including various multimodal task pipline (e.g., image captioning, VQA, video-text retrieval), customizable dataset (e.g., MS-COCO, ActivityNet, MSR-VTT), pre-trained model acquire (e.g., CLIP, BLIP-2)

2

McQuic. Repository of CVPR'22 paper "Unified Multivariate Gaussian Mixture for Efficient Neural Image Compression"

2

EMCL. [NeurIPS 2022] Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations

2

metrics. 📊 An infographics generator with 30+ plugins and 200+ options to display stats about your GitHub account and render them as SVG, Markdown, PDF or JSON!

2

rich. Rich is a Python library for rich text and beautiful formatting in the terminal.

2

LMaaS-Papers. Awesome papers on Language-Model-as-a-Service (LMaaS)

2

Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models

1

Awesome-Weak-to-Strong-Generalization. 🔥🔥🔥 This repository curates research on Weak-to-Strong Generalization across LLMs, multimodal learning, and beyond, focusing on how strong models learn from weak supervision and surpass their teachers. Stay tuned for the latest updates!

1
29
Apply