PhD student & MS at PKU, Intern @Tencent, Ex intern @Bytedance, @Microsoft Asia, @SenseTime-FVG, @megvii-research
Open-R1-Video. ✨First Open-Source R1-like Video-LLM [2025/02/18]
382Awesome-Multimodal-Large-Language-Models. 🔥Awesome Multimodal Large Language Models Paper List
154SSHT. ICASSP 2024 "Learning Invariant Representation with Consistency and Diversity for Semi-supervised Source Hypothesis Transfer"
17Long-DWM. 🌟[AAAI 2026] The official repo for "LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model"
12LeCo_UDA. ACCV 2022 "Revisiting Unsupervised Domain Adaptation Models: a Smoothness Perspective"
8GPT2-BaikeChat. Peking University (PKU-SS) 2022 AI Program (Excellent) : "Encyclopedia chatbot using GPT"
7LeanPO. [CVPR 2026 Findings] "LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs"
5DADM. Domian Adaptation via Diffusion Models
2Person_Detection_KalmanFilter. homework for digital image processing
2SSHT-plus-plus. SSHT++
2RL-V. A fork of RLAIF-V
2MCTS-DPO. Jupyter Notebook
2D3-Video-webpage.
1lmms-eval. Accelerating the development of large multimodal models (LMMs) with one-click evaluation module - lmms-eval.
1diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image and audio generation in PyTorch
1sdm. A latent text-to-image diffusion model
1LLaVA-NeXT. Python
1EGTR. Python
1PGP. Code for "Multimodal Trajectory Prediction Conditioned on Lane-Graph Traversals," CoRL 2021.
1Visual-ChatGPT. 34K Star⭐! Xiaodong Wang is a core author of "Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models"
1llava-onevision. Python
1Llama-X. Open Academic Research on Improving LLaMA to SOTA LLM
1NUWA--3D. Xiaodong Wang is the first author of IJCAI 2023 "NUWA-3D: Learning 3D photography videos via self-supervised diffusion on single images"
1Easy_Domain_Adaptation_Framework. This is a simple yet effective framework for Domain Adaptation.
1ms-swift-RL. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 500+ LLMs (Qwen3, Qwen3-MoE, Llama4, GLM4.5, InternLM3, DeepSeek-R1, ...) and 200+ MLLMs (Qwen2.5-VL, Qwen2.5-Omni, InternVL3.5, Ovis2.5, Llava, GLM4v, Phi4, ...) (AAAI 2025).
1