Princeton, NJ

Wenhao Chai

Elite
@wenhaochai

CS Ph.D. @ Princeton University

StableVideo. [ICCV 2023] StableVideo: Text-driven Consistency-aware Diffusion Video Editing

1.4k

MovieChat. [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

706

Awesome-VQVAE. A collection of resources and papers on Vector Quantized Variational Autoencoder (VQ-VAE) and its application

333

aurora. [ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

147

CityGen. 🏙️🌆🌃 Try Infinite and Controllable 3D City Layout Generation!

45

STEVE. [ECCV 2024] STEVE in Minecraft is for See and Think: Embodied Agent in Virtual Environment

41

Awesome-DriveLM. 📚 A collection of resources and papers on Large Language Models in autonomous driving

27

PoseDA. [ICCV 2023] Global Adaptation meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose Estimation

24

UIUC-CS357-22SP. Workspace for CS357

19

UniAP. [AAAI 2024] UniAP: Towards Universal Animal Perception in Vision via Few-shot Learning

12

Self-supervised-Cross-view-3D-Human-Pose-Estimation-and-Localization-in-Video. A algorithm to process 3D-multi cross-view dataset based on Human3.6M or others, and realize the mapping from 2D joints location to 3D in our dataset.

7

awesome-cvpr2022. workshop, tutorial, oral, and poster with notes in cvpr2022

6

project-page-template. CSS

5

Apollo. Apollo is a family of LMMs designed for video understanding

5

video-dataset-maker. A pipeline covers downloading videos from YouTube and extracting frames using ffmpeg.

4

Random-Bridge-Generator. a blender platform for developing computer vision-based structural inspection algorithms

3

old_web. personal website built on beautiful jekyll, feel free to clone and modify

3

Structural-Health-Monitoring-HRNet. Codes for the competition IC-SHM 2021.

3

UniVHP. Unified Human-centric Perception Model and Benchmark in Sports

2

Missing-Label-Detection. With imperfect bounding box annotation, 30% of missing labels in this project, normal detection method like YOLOv5 doesn’t achieve a relatively good result. In our project, we use COCO dataset. And we greatly eliminate the negative influence on missing labels by using a modified loss function and dynamic weight.

2

Awesome_Prompting_Papers_in_Computer_Vision. A curated list of prompt-based paper in computer vision and vision-language learning.

1

pose2img. pose-driven human natural image generation based on latent diffusion model

1

AI-Hackathon-Molecular-Dynamics. Codes for AI Hackathon Molecular Dynamics.

1
23
Apply