This is your work, valued
Researcher at the Beijing Academy of Artificial Intelligence (BAAI);Postdoc at Tsinghua University.
Point-BERT. [CVPR 2022] Pre-Training 3D Point Cloud Transformers with Masked Point Modeling
Common-envs-issues. Just to document the environmental pitfalls I've encountered while training large models.
SpaceR. SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoning
InternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
vggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
dust3r. DUSt3R: Geometric 3D Vision Made Easy
Awesome-LLM-3D. Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resources
Swin3D_Task. The Experiment Code for Swin3D
See3D. [CVPR'25 Highlight] You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
OmniGen. OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
LLM101n. LLM101n: Let's build a Storyteller
big_vision. Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more.
tokenize-anything. [ECCV 2024] Tokenize Anything via Prompting
3DTrans. An open-source codebase for exploring autonomous driving pre-training
RSPrompter. This is the pytorch implement of our paper "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model"
samexporter. Exporting Segment Anything, MobileSAM, and Segment Anything 2 into ONNX format for easy deployment
Matcher. [ICLR'24 & IJCV‘25] Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
OpenPCDet. OpenPCDet Toolbox for LiDAR-based 3D Object Detection.
MDE-SpikingCamera. [ECCV2022] Spike Transformer: Monocular Depth Estimation for Spiking Camera