CS PhD @ Johns Hopkins University Multimodal & Spatial reasoning
Two-Stage-Gan-in-trajectory-generation. A Two-Stage Gan architecture to generate trajectory conditioned on maps information.
48Animefy. A "selfie2anime" project based on StyleGAN & StyleGAN2. You can generate the customed animate faces base on your own real-world selfie.
263D-Aware-VQA. Official Code for the NeurIPS'23 paper "3D-Aware Visual Question Answering about Parts, Poses and Occlusions"
21Spatial457. [CVPR'25 Highlight] A VQA benchmark for 6D spatial reasoning.
20DynSuperCLEVR. A video question answering dataset that focuses on the dynamics properties of objects (velocity, acceleration) and their collisions within 4D scenes.
20wheres-waldo. A Pytorch implementing of A Deep Learning approach to Template Matching. Usie Hypernet + VGG to match the templates.
13ruc_traffic_prediction. 交通数据是典型的多周期时间序列,我通过MSARIMA模型和DS Holt-Winters指数平滑模型对数据的周期性信息进行拟合和提取,实现车流速度的预测。
11superclevr-3D-question. 3D-aware visual question answering dataset, for parts, poses and occlusion reasoning. Published in NeurIPS'23.
10KeyVID. Offical code of paper KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation.
6XModBench. XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
6DeepIV. Implementation of Deep IV: A Flexible Approach for Counterfactual Prediction by TensorFlow 2
6iv2sr. An R package reimplementing of Regularization Methods for High-Dimensional Instrumental Variables Regression With an Application to Genetical Genomics.
2Trajectory_process. A python project to about trajectory process
1