Boston, MA

Arijit Ray

Advanced
@arijitray1993

Computer Vision Ph.D. Student, Boston University

awesome-spatial-reasoning. Collection of the latest spatial, 3D, and video/temporal reasoning papers

36

SAT. Spatial Aptitude Training for Multimodal Langauge Models

33

COLA. COLA: Evaluate how well your vision-language model can Compose Objects Localized with Attributes!

25

CARLA_tutorial. Jupyter Notebook

21

VQARelevance. Models and Codes for the paper Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions

14

mull. Repository to train multimodal latent reasoning tokens for Qwen 2.5 VL.

12

music_video_gen. Music Video Generation using a Deep Generative Adversarial Network

4

ConVQA. Data and models for "Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation"

3

minerva. collaborative agent-human slides

3

socratis. A benchmark of diverse emotional reactions and explanations for image-caption pairs

2

raspberrypi_surveillance. Python

2

errormap_atteneval. Code for the paper: Pointing to Error-Inducing Regions to Improve Explanation Helpfulness

2

practical_tmux. Barebones practical tmux configuration with keyboard-based fast switching and resizing because mouse-mode sucks for copy-pasting

1

formal_methods1. Lean

1

rpi_alarm. WiFi-based alarm using Raspberry Pi

1

utils. Common utils I use for multiple projects

1

helpfulness_evaluation. Evaluate helpfulness of heatmap-based explanations for predicting model performance

1

arijitray1993.github.io. Personal Webpage

1

multimodal_thinking. HTML

1

TensorFlowRBF. This repository explores the design of a Radial Basis Function and related functions (like K-Means) for use with TensorFlow.

1
20
Apply