Vision-and-Language-Benchmark. Codebase for research of vision&language, including various multimodal task pipline (e.g., image captioning, VQA, video-text retrieval), customizable dataset (e.g., MS-COCO, ActivityNet, MSR-VTT), pre-trained model acquire (e.g., CLIP, BLIP-2)

github.com/zchoi/Vision-and-Language-Benchmark

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.