Computer Vision Ph.D. Student, Boston University
awesome-spatial-reasoning. Collection of the latest spatial, 3D, and video/temporal reasoning papers
36SAT. Spatial Aptitude Training for Multimodal Langauge Models
33COLA. COLA: Evaluate how well your vision-language model can Compose Objects Localized with Attributes!
25CARLA_tutorial. Jupyter Notebook
21VQARelevance. Models and Codes for the paper Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions
14mull. Repository to train multimodal latent reasoning tokens for Qwen 2.5 VL.
12music_video_gen. Music Video Generation using a Deep Generative Adversarial Network
4ConVQA. Data and models for "Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation"
3minerva. collaborative agent-human slides
3socratis. A benchmark of diverse emotional reactions and explanations for image-caption pairs
2raspberrypi_surveillance. Python
2errormap_atteneval. Code for the paper: Pointing to Error-Inducing Regions to Improve Explanation Helpfulness
2practical_tmux. Barebones practical tmux configuration with keyboard-based fast switching and resizing because mouse-mode sucks for copy-pasting
1formal_methods1. Lean
1rpi_alarm. WiFi-based alarm using Raspberry Pi
1utils. Common utils I use for multiple projects
1helpfulness_evaluation. Evaluate helpfulness of heatmap-based explanations for predicting model performance
1arijitray1993.github.io. Personal Webpage
1multimodal_thinking. HTML
1TensorFlowRBF. This repository explores the design of a Radial Basis Function and related functions (like K-Means) for use with TensorFlow.
1