A Ph.D. student in KAUST. My current focus is vision-language learning.
ViECap. Transferable Decoding with Visual Entities for Zero-Shot Image Captioning, ICCV 2023
167Tempo. Tempo: Small Vision-Language Models are Smart Compressors for Long Video Understanding, ECCV 2026
78awesome-zero-shot-captioning. A curated list of zero-shot captioning papers
24feielysia.github.io. This is the homepage of Junjie Fei
1