Senior Research Scientist at TikTok. I got my Ph.D from University of Chinese Academy of Sciences (UCAS), with a research focus on CV, LLM, AI Agent.
yolov12. [NeurIPS 2025] YOLOv12: Attention-Centric Real-Time Object Detectors
2.9kiTPN. (CVPR2023/TPAMI2024) Integrally Pre-Trained Transformer Pyramid Networks -- A Hierarchical Vision Transformer for Masked Image Modeling
216SDL-Skeleton. A toolbox for object skeleton detection, can also be used for edge detection, building extraction and road extraction. TIP (2021)
138ChatterBox. [AAAI2025] ChatterBox: Multi-round Multimodal Referring and Grounding, Multimodal, Multi-round dialogues
61beyond_masking. Beyond Masking: Demystifying Token-Based Pre-Training for Vision Transformers
26SaGe. (SaGe) Semantic-Aware Generation for Self-Supervised Visual Representation Learning
26GOLD_NAS. The implementation of GOLD_NAS
24DAAS. 'Discretization-Aware Architecture Search' alleviates the discretization gap in one-shot differentiable NAS. DAAS has been accepted by PR (2021).
20IFNAS. Python
1yolov10. YOLOv10: Real-Time End-to-End Object Detection [NeurIPS 2024]
1