4MC-4M-Image-Text-Pairs-with-CLIP-embeddings. I have created a dataset of Image-Text-Pairs by using the cosine similarity of the CLIP embeddings of the image & it's caption derrived from YFCC100M. I have also added propabilities from a NSFW detector & more.

github.com/christophschuhmann/4MC-4M-Image-Text-Pairs-with-CLIP-embeddings

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.