Xiaodong Wang

Expert
@Wang-Xiaodong1899

PhD student & MS at PKU, Intern @Tencent, Ex intern @Bytedance, @Microsoft Asia, @SenseTime-FVG, @megvii-research

Open-R1-Video. ✨First Open-Source R1-like Video-LLM [2025/02/18]

382

Awesome-Multimodal-Large-Language-Models. 🔥Awesome Multimodal Large Language Models Paper List

154

SSHT. ICASSP 2024 "Learning Invariant Representation with Consistency and Diversity for Semi-supervised Source Hypothesis Transfer"

17

Long-DWM. 🌟[AAAI 2026] The official repo for "LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model"

12

LeCo_UDA. ACCV 2022 "Revisiting Unsupervised Domain Adaptation Models: a Smoothness Perspective"

8

GPT2-BaikeChat. Peking University (PKU-SS) 2022 AI Program (Excellent) : "Encyclopedia chatbot using GPT"

7

LeanPO. [CVPR 2026 Findings] "LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs"

5

DADM. Domian Adaptation via Diffusion Models

2

Person_Detection_KalmanFilter. homework for digital image processing

2

SSHT-plus-plus. SSHT++

2

RL-V. A fork of RLAIF-V

2

MCTS-DPO. Jupyter Notebook

2

D3-Video-webpage.

1

lmms-eval. Accelerating the development of large multimodal models (LMMs) with one-click evaluation module - lmms-eval.

1

diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image and audio generation in PyTorch

1

sdm. A latent text-to-image diffusion model

1

LLaVA-NeXT. Python

1

EGTR. Python

1

PGP. Code for "Multimodal Trajectory Prediction Conditioned on Lane-Graph Traversals," CoRL 2021.

1

Visual-ChatGPT. 34K Star⭐! Xiaodong Wang is a core author of "Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models"

1

llava-onevision. Python

1

Llama-X. Open Academic Research on Improving LLaMA to SOTA LLM

1

NUWA--3D. Xiaodong Wang is the first author of IJCAI 2023 "NUWA-3D: Learning 3D photography videos via self-supervised diffusion on single images"

1

Easy_Domain_Adaptation_Framework. This is a simple yet effective framework for Domain Adaptation.

1

ms-swift-RL. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 500+ LLMs (Qwen3, Qwen3-MoE, Llama4, GLM4.5, InternLM3, DeepSeek-R1, ...) and 200+ MLLMs (Qwen2.5-VL, Qwen2.5-Omni, InternVL3.5, Ovis2.5, Llava, GLM4v, Phi4, ...) (AAAI 2025).

1
25
Apply