This is your work, valued
PhD student in University of Trento, Italy.
Detect-and-read-meters. This is the first released system towards complex meters` detection and recognition, which is implemented by computer vision techniques.
★ 196Earth-Observation-VLMs. 🔥🔥A Family of Multi-Sensor, Multi-Granularity Vision-Language Models for Earth Observation Understanding
★ 137Visual-Text-Processing-survey. The official project of paper "Visual Text Processing: A Comprehensive Review and Unified Evaluation""
★ 103MLLM-Semantic-Hallucination. 🔥🔥[NeurIPS2025]Exploring and mitigating semantic hallucinations in scene text perception and reasoning
★ 30VidText. Comprehensive benchmark for video text understanding
★ 29Efficient-Ambiguous-Text-Detector. An official Project related to Paper "Perceiving Ambiguity and Semantics without Recognition: An Efficient and Effective Ambiguous Scene Text Detector" (ACM MM 2023)
★ 22Pattern-Recognition-Algorithm. Re-implementation of some classical algorithm in pattern recognition
★ 3scripts-for-image-processing. some practical demos for image/ text processing
★ 1Synthesis-multilingual-handwritten-text-data. This is a simple yet method focused on handwritten text dataset generation, which is beneficial for handwritten text detection and segmentation
★ 1multilingual-machine-translation. This is some code for multilingual machine translation (English, Korean, Japanese, Arabic)
★ 1TextBorder. official implementation for paper 《BAG:Learning Border Attraction Grouping for Arbitrary-shaped Scene Text Detection》
★ 1meeting_reports. some presentations for weekly meetings
★ 1terrascope. HTML
★ 1DepthART. [ACMMM 2026] Official code for DepthART: Depth Anything Rethought for Tiny Models
★ 98VideoCreator. TypeScript
★ 1DyCo-RL. DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
★ 18unireason-med. Python
★ 5lightly-studio. Curate, Annotate, and Manage Your Data in LightlyStudio.
★ 869PoInit-of-View. [CVPR-26🎉🎉🎉] This is the official repository for "PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems".
★ 12coding-with-beat. 🎵 将音乐搬进AI终端 · 打造Coding专属智能DJ · 听歌新范式 · 交互式音乐 |A retro pixel DJ for Claude Code / Codex CLI / Terminal with Apple Music / QQ Music / Local Music — plays music, shows synced lyrics, and 😱 panics when your tests fail.
★ 114UNEM-Transductive. (CVPR 2025) Official code of the paper: UNEM: UNrolled Generalized EM for Transductive Few-Shot Learning
★ 13AdvFLYP. [CVPR-26 finding 🎉] This is the official repository for our work "Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models".
★ 27EAGC. (CVPR2026 Highlight) The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery (EAGC)
★ 33SEM. [CVPR Findings 2026] SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
★ 21MemCoach. [CVPR'26 Highlight] MemCoach: Steering-based MLLM for Actionable Image Memorability Feedback
★ 42proactivebench. Official repository of "ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models" (ECCV 2026)
★ 29CAD. [ECCV 2026] Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
★ 28Video-Browser. Official code repo of Video-Browser: Towards Agentic Open-web Video Browsing
★ 28daily_stock_analysis. LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
★ 60kVideo-Next-Event-Prediction. Python
★ 28MLLM-Semantic-Hallucination. 🔥🔥[NeurIPS2025]Exploring and mitigating semantic hallucinations in scene text perception and reasoning
★ 30RePro. The official code of Refinement Provenance Inference: Detecting LLM-Refined Training Prompts from Model Behavior
★ 22IDO. Turn every moment into momentum
★ 22OWDFA-CAL. (AAAI2026) Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning (CAL)
★ 32wlp. Loomis Painter: Reconstructing the painting process
★ 55Fleming-VL. Fleming-VL: Towards Universal Medical Visual Understanding with Multimodal LLMs
★ 15verl-internvl. Python
★ 53Fleming-R1. Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
★ 31VideoX22L. Python
★ 6Awesome-Medical-Dataset. Collection of awesome medical dataset resources.
★ 2.1kAwesome_Think_With_Images. Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
★ 1.5kMed-VLM-Bench-Summary. A Curated Benchmark Repository for Medical Vision-Language Models
★ 198GlyphOnly. 【2024 ECAI】First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
★ 14Earth-Observation-VLMs. 🔥🔥A Family of Multi-Sensor, Multi-Granularity Vision-Language Models for Earth Observation Understanding
★ 137VidText. Comprehensive benchmark for video text understanding
★ 29Awesome-Multimodal-Large-Language-Models. 🔥Awesome Multimodal Large Language Models Paper List
★ 154CoS_codes. CoS: Chain-of-Shot Prompting for Long Video Understanding
★ 53Sa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.7kmglmm. Python
★ 32MegaPairs. [ACL 2025 Oral] 🔥🔥 MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval
★ 248SAMRS. The official repo for [NeurIPS'23] "SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model"
★ 385OMG-Seg. Official Repo For OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]
★ 1.4kIEEE_TPAMI_SpectralGPT. Hong, D., Zhang, B., Li, X., Li, Y., Li, C., Yao, J., Yokoya, N., Li, H., Ghamisi, P., Jia, X., Plaza, A. and Gamba, P., Benediktsson, J., Chanussot, J. (2024). SpectralGPT: Spectral remote sensing foundation model. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. DOI:10.1109/TPAMI.2024.3362475.
★ 278VoCo-LLaMA. [CVPR'2025] VoCo-LLaMA: This repo is the official implementation of "VoCo-LLaMA: Towards Vision Compression with Large Language Models".
★ 205Video-XL. 🔥🔥First-ever hour scale video understanding models
★ 626VideoNIAH. VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMs
★ 57TextCtrl. [2024-NeurIPS] TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
★ 106OmniGen. OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
★ 4.3kOpen-LLaVA-NeXT. An open-source implementation for training LLaVA-NeXT.
★ 439VILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kVISTA_Evaluation_FineTuning. Evaluation code and datasets for the ACL 2024 paper, VISTA: Visualized Text Embedding for Universal Multi-Modal Retrieval. The original code and model can be accessed at FlagEmbedding.
★ 48LongVA. Long Context Transfer from Language to Vision
★ 407SpatialBot. The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.
★ 349MLVU. 🔥🔥MLVU: Multi-task Long Video Understanding Benchmark
★ 266MMVU. Python
★ 57FlagEmbedding. Retrieval and Retrieval-augmented LLMs
★ 12kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kMovieChat. [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
★ 706Bunny. A family of lightweight multimodal models.
★ 1.1kAwesome_Long_Form_Video_Understanding. Awesome papers & datasets specifically focused on long-term videos.
★ 381experimental-consistory. Python
★ 113VideoRecap. Python
★ 209SmartEdit. Official code of SmartEdit [CVPR-2024 Highlight]
★ 374ICV. Code for In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
★ 202Efficient-Ambiguous-Text-Detector. An official Project related to Paper "Perceiving Ambiguity and Semantics without Recognition: An Efficient and Effective Ambiguous Scene Text Detector" (ACM MM 2023)
★ 22Visual-Text-Processing-survey. The official project of paper "Visual Text Processing: A Comprehensive Review and Unified Evaluation""
★ 103time-diffusion. Official code repo for "Editing Implicit Assumptions in Text-to-Image Diffusion Models"
★ 89Recommendations-Diffusion-Text-Image. A paper collection of recent diffusion models for text-image generation tasks, e,g., visual text generation, font generation, text removal, text image super resolution, text editing, handwritten generation, scene text recognition and scene text detection.
★ 273bolei_awesome_posters. CVPR and NeurIPS poster examples and templates
★ 2ksentence-transformers. State-of-the-Art Embeddings, Retrieval, and Reranking
★ 19kOCR-TR. Optocal Character Recognition (OCR / HTR) using Transformers
★ 11multilingual-machine-translation. This is some code for multilingual machine translation (English, Korean, Japanese, Arabic)
★ 1SSL-OCR. Text-DIAE: A Self-Supervised Degradation Invariant Autoencoders for Text Recognition and Document Enhancement - AAAI 2023
★ 30Synthesis-multilingual-handwritten-text-data. This is a simple yet method focused on handwritten text dataset generation, which is beneficial for handwritten text detection and segmentation
★ 1lama. 🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022
★ 10kOCR-SAM. [Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting
★ 590Prompt-Segment-Anything. This is an implementation of zero-shot instance segmentation using Segment Anything.
★ 316Detect-and-read-meters. This is the first released system towards complex meters` detection and recognition, which is implemented by computer vision techniques.
★ 197Medical-SAM-Adapter. A lightweight adapter bridges SAM with medical imaging [MedIA]
★ 1.3kfinetune-anything. Fine-tune SAM (Segment Anything Model) for computer vision tasks such as semantic segmentation, matting, detection ... in specific scenarios
★ 868peft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kpaper_downloader. Download papers and supplemental materials from open-access paper website, such as AAAI, AAMAS, AISTATS, COLT, CORL, CVPR, ECCV, ICCV, ICLR, ICML, IJCAI, JMLR, NIPS, RSS, WACV.
★ 305singa. a distributed deep learning platform
★ 3.6kccf-deadlines. ⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
★ 9.2kOCR_DataSet. 收集并整理有关OCR的数据集并统一标注格式,以便实验需要
★ 970WXData. Python
★ 1ICDAR2COCO. A tool for the conversion from ICDAR to COCO dataset.
★ 9OpenAI-CLIP. Simple implementation of OpenAI CLIP model in PyTorch.
★ 725meeting_reports. some presentations for weekly meetings
★ 1datasets. 收藏一些数据集及其下载脚本
★ 2TextBorder. official implementation for paper 《BAG:Learning Border Attraction Grouping for Arbitrary-shaped Scene Text Detection》
★ 1FeedBack. Kaggle Feedback Prize - Evaluating Student Writing 15th solution
★ 8TRM_tutorial. Transformer在CV和NLP领域的变体模型的从零解读:Transformer;VIT;Swin Transformer
★ 337External-Attention-pytorch. 🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐
★ 3External-Attention-pytorch. 🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐
★ 12kPattern-Recognition-Algorithm. Re-implementation of some classical algorithm in pattern recognition
★ 3alpr_utils. ALPR model in unconstrained scenarios for Chinese license plates
★ 173CRAFT-Reimplementation. CRAFT-Pyotorch:Character Region Awareness for Text Detection Reimplementation for Pytorch
★ 467