Ph.D @ University of Washington | AI Researcher, Adventurer, Philosopher.
video-to-audio-through-text. [NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”
38xRIR_code. [CVPR 2025] Pytorch implementation of the paper "Hearing Anywhere in Any Environment"
34Vision-to-Audio-and-Beyond. ICML 2024 "From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation"
10CVPR-2024-Speech_Audio_Music-Papers. A curated collections of papers related to speech, audio and music in CVPR 2024.
7MUSIC-AVQA-v2.0. Additional Videos Data and QA pairs for Balancing Original MUSIC-AVQA Dataset
6BLDC-Sensorless-Motor-Control-using-Microchip-embedded-system.
2hearinganywhereinanyenvironment. Homepage for the work "Hearing Anywhere in Any Environment" (CVPR 2025)
1CSE546-Project. Character-level language models and sequence modelling
1old-website. Personal website
1