Working on 3D perception, vision and language
MSMDFusion. [CVPR 2023] MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object Detection
207UniToken. [CVPRW 2025] UniToken is an auto-regressive generation model that combines discrete and continuous representations to process visual inputs, making it easy to integrate both visual understanding and image generation tasks seamlessly.
106Lumen. [NeurIPS 2024] Lumen: a Large multimodal model with versatile vision-centric capabilities
25Transformer-backbone. The reproduce of Transformer architecture in paper "Attention is all your need"
18MORE. [ECCV 2022] MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes official implementation
16automatic-matting. The project aims to extract portrait from a picture automatically
5TV-Net. The official code of MM 2021 paper "Two-stage Visual Cues Enhancement Network for Referring Image Segmentation"
3UESTC_OS_experiment. 电子科技大学操作系统进程与资源管理实验代码(python)
3mcu_design. 三天三夜小组mcu设计
3MDU_preprocess. Creating virtual 2D points via multi-depth projection
1