clip-synthetic-captions. Tiny-scale experiment showing that CLIP models trained using detailed captions generated by multimodal models (CogVLM and LLaVA 1.5) outperform models trained using the original alt-texts on a range of classification and retrieval tasks.

github.com/nopperl/clip-synthetic-captions

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.