CS Ph.D. @ Princeton University
StableVideo. [ICCV 2023] StableVideo: Text-driven Consistency-aware Diffusion Video Editing
1.4kMovieChat. [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
706Awesome-VQVAE. A collection of resources and papers on Vector Quantized Variational Autoencoder (VQ-VAE) and its application
333aurora. [ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
147CityGen. 🏙️🌆🌃 Try Infinite and Controllable 3D City Layout Generation!
45STEVE. [ECCV 2024] STEVE in Minecraft is for See and Think: Embodied Agent in Virtual Environment
41Awesome-DriveLM. 📚 A collection of resources and papers on Large Language Models in autonomous driving
27PoseDA. [ICCV 2023] Global Adaptation meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose Estimation
24UIUC-CS357-22SP. Workspace for CS357
19UniAP. [AAAI 2024] UniAP: Towards Universal Animal Perception in Vision via Few-shot Learning
12Self-supervised-Cross-view-3D-Human-Pose-Estimation-and-Localization-in-Video. A algorithm to process 3D-multi cross-view dataset based on Human3.6M or others, and realize the mapping from 2D joints location to 3D in our dataset.
7awesome-cvpr2022. workshop, tutorial, oral, and poster with notes in cvpr2022
6project-page-template. CSS
5Apollo. Apollo is a family of LMMs designed for video understanding
5video-dataset-maker. A pipeline covers downloading videos from YouTube and extracting frames using ffmpeg.
4Random-Bridge-Generator. a blender platform for developing computer vision-based structural inspection algorithms
3old_web. personal website built on beautiful jekyll, feel free to clone and modify
3Structural-Health-Monitoring-HRNet. Codes for the competition IC-SHM 2021.
3UniVHP. Unified Human-centric Perception Model and Benchmark in Sports
2Missing-Label-Detection. With imperfect bounding box annotation, 30% of missing labels in this project, normal detection method like YOLOv5 doesn’t achieve a relatively good result. In our project, we use COCO dataset. And we greatly eliminate the negative influence on missing labels by using a modified loss function and dynamic weight.
2Awesome_Prompting_Papers_in_Computer_Vision. A curated list of prompt-based paper in computer vision and vision-language learning.
1pose2img. pose-driven human natural image generation based on latent diffusion model
1AI-Hackathon-Molecular-Dynamics. Codes for AI Hackathon Molecular Dynamics.
1