This is your work, valued
FavoritePapers.
★ 460chainer-partial_convolution_image_inpainting. Reproduction of Nvidia image inpainting paper "Image Inpainting for Irregular Holes Using Partial Convolutions"
★ 114CLIP-visualization. Attention visualization in CLIP
★ 17summary-to-pptx. paper slide summary maker using python-pptx on Windows
★ 10MTRNN. implementation of multiple timescale recurrent neural network
★ 8simple_beamsearch. simple_beamsearch example
★ 5chainer-StarGAN. chainer implementation of StarGAN
★ 4bibtex-abbreviator. Convert bibtex into shorter style
★ 3DCGAN-chainer. DCGAN implementation using chainer
★ 3VandL-Walking-Guide. Vision and Language (V&L) 研究の歩き方
★ 1hermes-agent. The agent that grows with you
★ 223kpdfannots. Extracts and formats text annotations from a PDF file
★ 669pySBD. 🐍💯pySBD (Python Sentence Boundary Disambiguation) is a rule-based sentence boundary detection that works out-of-the-box.
★ 927karukan. Japanese Input Method System for Linux, macOS, Neural Kana-Kanji Conversion Engine
★ 688academic-research-skills. Academic Research Skills for Claude Code: research → write → review → revise → finalize
★ 40kclaude-session-tracker. Auto-save every Claude Code session to GitHub Projects. Never lose Prompt & Context again.
★ 46AutoResearchClaw. Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
★ 14kcopilot-cli-for-beginners. Learn how to get started using the GitHub Copilot CLI!
★ 2.8kccusage. npx ccusage
★ 18kECC. The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
★ 236kpm-skills. PM Skills Marketplace: 100+ agentic skills, commands, and plugins — from discovery to strategy, execution, launch, and growth.
★ 25kRepresentation-as-a-judge. The code for paper "Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry", accepted by ICLR 2026.
★ 218SE-Bench. Official repo for "SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization"
★ 28metamorph. Code for MetaMorph Multimodal Understanding and Generation via Instruction Tuning
★ 235embedding-atlas. Embedding Atlas is a tool that provides interactive visualizations for large embeddings. It allows you to visualize, cross-filter, and search embeddings and metadata.
★ 4.9kBabyVision. We introduce BabyVision, a benchmark revealing the infancy of AI vision.
★ 235evalscope. A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
★ 3.2kSurvey4MusicAVQA. Survey for MusicAVQA
★ 5JudgeAnything. Python
★ 17Daily-Omni. This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
★ 46AI-Agents-Projects-Tutorials. Multi-agent systems, memory, planning, reasoning loops
★ 2.8kpytest-ml-tdd-example. 「テストを書かない研究者に送る “最初にテストを書く” 実験コード入門 – オレオレ最強 main.py から抜け出すために –」 のサンプルコード
★ 9live-vlm-webui. Real-time Vision Language Model interaction via webcam - WebRTC-based web interface
★ 408awesome-hallucination-detection. List of papers on hallucination detection in LLMs.
★ 1.1kDeepSeek-OCR. Contexts Optical Compression
★ 24kawesome-vl-compositionality. Awesome Vision-Language Compositionality, a comprehensive curation of research papers in literature.
★ 40DACoN. Python
★ 21BasicPBC. Official Implementation of "Learning Inclusion Matching for Animation Paint Bucket Colorization"
★ 306RLAIF-DialogLLM. Python
★ 7VibeCheck. Automated Qualitative Analysis of LLMs (ICLR 2025)
★ 53harmony. Renderer for the harmony response format to be used with gpt-oss
★ 4.5kEvaLearn. EvaLearn is a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks.
★ 431Awesome-LLM-as-a-judge.
★ 570MiCo. [ICCV 2025] Explore the Limits of Omni-modal Pretraining at Scale
★ 124ONLY. [ICCV 2025] ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
★ 51UPD. [ACL2025] Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
★ 82MergeToVLRM. Source code of our paper "Transferring Textual Preferences to Vision-Language Understanding through Model Merging", ACL 2025
★ 5MLLM-Judge. [ICML 2024 Oral] Official code repository for MLLM-as-a-Judge.
★ 95llm-jp-judge. 生成自動評価を行うためのPythonツール
★ 54Awesome_Think_With_Images. Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
★ 1.5ktorchtitan. A PyTorch native platform for training generative AI models
★ 5.6ksrunx. A modern Python library for SLURM workload manager integration with workflow orchestration capabilities.
★ 16Awesome-Unified-Multimodal-Models. 📖 This is a repository for organizing papers, codes and other resources related to unified multimodal models.
★ 829Awesome-Unified-Multimodal-Models. Awesome Unified Multimodal Models
★ 1.3kFinMME. [ACL 2025] FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
★ 68policy. TypeScript
★ 398Evaluation-Agent. [ACL2025 Oral & Award] Evaluate Image/Video Generation like Humans - Fast, Explainable, Flexible
★ 128TRACT. Python
★ 24vision-explanation-methods. Methods for creating saliency maps for computer vision models.
★ 45pixmo-docs. ACL 2025: Synthetic data generation pipelines for text-rich images.
★ 168awesome_cs-ja_phd_life. collection of articles about PhD life written in 🇯🇵
★ 342visual_inference_chain. This repository contains the official code for our paper: Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
★ 25gpt-cost-estimator. A cost estimator for OpenAI API calls in tqdm loops.
★ 20gpu-aquarium. JavaScript
★ 58AutoConverter. Official implementation of "Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation" (CVPR 2025)
★ 40llm-jp-eval-mm. A lightweight framework for evaluating visual-language models.
★ 43Awesome-LVLM-Hallucination. up-to-date curated list of state-of-the-art Large vision language models hallucinations research work, papers & resources
★ 325From-Redundancy-to-Relevance. [NAACL 2025 Oral] From redundancy to relevance: Enhancing explainability in multimodal large language models
★ 130web-ui. 🖥️ Run AI Agent in your browser.
★ 16kSlideVQA. SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images (AAAI2023)
★ 106Open-LLM-Leaderboard. Open-LLM-Leaderboard: Open-Style Question Evaluation. Paper at https://arxiv.org/abs/2406.07545
★ 53Sa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.7kVision-Language-Models-Overview. A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
★ 684japanese-clip-qwen2_vl. Jupyter Notebook
★ 2DINO-X-API. DINO-X: The World's Top-Performing Vision Model for Open-World Object Detection and Understanding
★ 1.4kpfgen-bench. Preferred Generation Benchmark
★ 103LLM-generated-Text-Detection. A survey and reflection on the latest research breakthroughs in LLM-generated Text detection, including data, detectors, metrics, current issues and future directions.
★ 249clipscore. CLIPScore EMNLP code
★ 251vllm-safety-benchmark. [ECCV 2024] Official PyTorch Implementation of "How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs"
★ 90Awesome-LVLM-Attack. 😎 up-to-date & curated list of awesome Attacks on Large-Vision-Language-Models papers, methods & resources.
★ 566Awesome-MLLM-Safety. Accepted by IJCAI-24 Survey Track
★ 233Visual-Adversarial-Examples-Jailbreak-Large-Language-Models. Repository for the Paper (AAAI 2024, Oral) --- Visual Adversarial Examples Jailbreak Large Language Models
★ 282HarmBench. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
★ 1kPrism. A Framework for Decoupling and Assessing the Capabilities of VLMs
★ 44FigStep. [AAAI'25 (Oral)] Jailbreaking Large Vision-language Models via Typographic Visual Prompts
★ 211VLGuard. [ICML 2024] Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.
★ 90MLLMGuard. Python
★ 4616824-CaLORAify. Python
★ 50Inpaint-Anything. Inpaint anything using Segment Anything and inpainting models.
★ 7.7kfiftyone. Refine high-quality datasets and visual AI models
★ 11kMulti-modal-Self-instruct. The codebase for our EMNLP24 paper: Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
★ 85ImgTrojan. Code and data for "ImgTrojan: Jailbreaking Vision-Language Models with ONE Image"
★ 24gRefCOCO. A benchmark dataset for GREx: GRES, GREC, and GREG [CVPR 2023 & IJCV 2026]
★ 241FaithD2T. Dataset and Code for Generating Faithful and Salient Text from Multimodal Data (INLG 2024) paper.
★ 1OmniDocBench. [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation
★ 1.9kflair. [CVPR 2025] FLAIR: VLM with Fine-grained Language-informed Image Representations
★ 148PathWeave. Code for paper "LLMs Can Evolve Continually on Modality for X-Modal Reasoning" NeurIPS2024
★ 41OmniBench. A project for tri-modal LLM benchmarking and instruction tuning.
★ 61Efficient-Multimodal-LLMs-Survey. Efficient Multimodal Large Language Models: A Survey
★ 387vstar. PyTorch Implementation of "V* : Guided Visual Search as a Core Mechanism in Multimodal LLMs"
★ 707GraphVis. Python
★ 7VITA. ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
★ 2.5kawesome-japanese-llm. 日本語LLMまとめ - Overview of Japanese LLMs
★ 1.4kALM-Bench. [CVPR 2025 🔥] ALM-Bench is a multilingual multi-modal diverse cultural benchmark for 100 languages across 19 categories. It assesses the next generation of LMMs on cultural inclusitivity.
★ 47semantic-entropy-probes. Jupyter Notebook
★ 65yomitoku. YomiTokuはAIを活用した日本語文書解析エンジンを提供するPythonパッケージです。 Yomitoku is an AI-powered document image analysis package designed specifically for the Japanese language.
★ 1.6kml-aim. This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.
★ 1.4kMMGenBench. Official repository of MMGenBench
★ 119mPLUG-DocOwl. mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2.4kevaluation-guidebook. Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
★ 2.1kAutoSurvey. Python
★ 471litellm. The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
★ 55kLRP-eXplains-Transformers. Layer-wise Relevance Propagation for Large Language Models and Vision Transformers [ICML 2024]
★ 243DocLayout-YOLO. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
★ 2.2krbo. Implementation of Rank-biased Overlap
★ 155docling. Get your documents ready for gen AI
★ 64kzerox. OCR & Document Extraction using vision models
★ 12kllama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
★ 19kWindowsAgentArena. Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
★ 884OmniParser. A simple screen parsing tool towards pure vision based GUI agent
★ 25kvlm-recipes. Python
★ 20MATRIX-Gen.
★ 47hypvl. This repository is related to 'Intriguing Properties of Hyperbolic Embeddings in Vision-Language Models', published at TMLR (2024), https://openreview.net/pdf?id=P5D2gfi4Gg
★ 21Awesome-Interpretability-in-Large-Language-Models. This repository collects all relevant resources about interpretability in LLMs
★ 402AwesomeLLM4SE. [SCIS 2025] A Survey on Large Language Models for Software Engineering
★ 338WorkBench. WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting. COLM 2024.
★ 73mid.metric. Python
★ 30Polos. [CVPR24 Highlights] Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
★ 33MMed-RAG. [ICLR'25] MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
★ 337donut. Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
★ 6.9ksynthtiger. Official Implementation of SynthTIGER (Synthetic Text Image Generator), ICDAR 2021
★ 579awesome-vlm-architectures. Famous Vision Language Models and Their Architectures
★ 1.3kswarm. Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
★ 22kOvis. A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
★ 1.5ksynth_doc_generation. Official PyTorch Implementation of DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis - ICDAR 2021
★ 93ChatDev. ChatDev 2.0: Dev All through LLM-powered Multi-Agent Collaboration
★ 34kMarigold. [CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
★ 3.2kagentscope. Build and run agents you can see, understand and trust.
★ 28kOpenreview. data from ICLR OpenReview and code for data analysis
★ 74haloscope. source code for NeurIPS'24 paper "HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection"
★ 70GPTGeoChat. Repository for "Granular Privacy Control for Geolocation with Vision Language Models"
★ 10LVCD. The official code of paper "LVCD: Reference-based Lineart Video Colorization with Diffusion Models"
★ 200Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kMMVP. Python
★ 365lmdeploy. LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
★ 8kvlm-evaluation. VLM Evaluation: Benchmark for VLMs, spanning text generation tasks from VQA to Captioning
★ 139AI-Scientist. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑🔬
★ 14kddpm-segmentation. Label-Efficient Semantic Segmentation with Diffusion Models (ICLR'2022)
★ 718Diffusion-based-Segmentation. This is the official Pytorch implementation of the paper "Diffusion Models for Implicit Image Segmentation Ensembles".
★ 318Teacher-free-Knowledge-Distillation. Knowledge Distillation: CVPR2020 Oral, Revisiting Knowledge Distillation via Label Smoothing Regularization
★ 582DriveAGI. [CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving
★ 802MetaCLIP. NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024
★ 1.9kMM-Vet. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities (ICML 2024)
★ 331GVIL. Code and data for EMNLP 2023 paper "Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?"
★ 15reflex. 🕸️ Web apps in pure Python 🐍
★ 29ksurya. OCR, layout analysis, reading order, table recognition in 90+ languages
★ 21ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kflexeval. Flexible evaluation tool for language models
★ 61evolutionary-model-merge. Official repository of Evolutionary Optimization of Model Merging Recipes
★ 1.4kAMEX-codebase. Python
★ 33BiVLC. Official repo for "BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval." NeurIPS 2024
★ 7COLA. COLA: Evaluate how well your vision-language model can Compose Objects Localized with Attributes!
★ 25Enhance-FineGrained. [CVPR 2024] Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Fine-grained Understanding
★ 56vision-language-models-are-bows. Experiments and data for the paper "When and why vision-language models behave like bags-of-words, and what to do about it?" Oral @ ICLR 2023
★ 294why-winoground-hard. Code for 'Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality', EMNLP 2022
★ 31Winoground-T2I.
★ 7VBench. [CVPR2024 Highlight] VBench - We Evaluate Video Generation
★ 1.7kEXAMS-V. A Multi-discipline Multilingual Multimodal Exam Benchmark
★ 5web2code. Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
★ 103MMT-Bench. [ICML 2024] | MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
★ 119VIMABench. Official Task Suite Implementation of ICML'23 Paper "VIMA: General Robot Manipulation with Multimodal Prompts"
★ 327MM-SafetyBench. Accepted by ECCV 2024
★ 218MuirBench. A Comprehensive Benchmark for Robust Multi-image Understanding
★ 21ml-ferret. Python
★ 8.7kINS-MMBench. [ICCV '25] INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance
★ 12LRP-for-ResNet. [ECCV24] Layer-Wise Relevance Propagation with Conservation Property for ResNet
★ 15ferret. A python package for benchmarking interpretability techniques on Transformers.
★ 214MLLM-Bench. MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
★ 78BEAF. [ECCV’24] Official repository for "BEAF: Observing Before-AFter Changes to Evaluate Hallucination in Vision-language Models"
★ 22FoolyourVLLMs. [ICML 2024] Fool Your (Vision and) Language Model With Embarrassingly Simple Permutations
★ 15vhs_benchmark. 🔥 [ICLR 2025] Official Benchmark Toolkits for "Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark"
★ 44Hemm. A holistic evaluation library for multi-modal generative models using Weave
★ 27DeCLIP. Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
★ 677Reflection_Tuning. [ACL'24] Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
★ 367FastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kWildVision-Bench. Python
★ 17TFC-pretraining. Self-supervised contrastive learning for time series via time-frequency consistency
★ 527yolov10. YOLOv10: Real-Time End-to-End Object Detection [NeurIPS 2024]
★ 11kyolov9. Implementation of paper - YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information
★ 9.5kTF-ID. TF-ID: Table/Figure IDentifier for academic papers
★ 247Transformer-MM-Explainability. [ICCV 2021- Oral] Official PyTorch implementation for Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers, a novel method to visualize any Transformer-based network. Including examples for DETR, VQA.
★ 911CLIP_Surgery. [Pattern Recognition 25] CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks
★ 483CLIP_Explainability. code for studying OpenAI's CLIP explainability
★ 39IllusionVQA. This repository contains the data and code of the paper titled "IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models"
★ 24iglu-datasets. Python
★ 43MileBench. This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"
★ 38VLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3kdata-juicer. Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.8kMMBench. Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"
★ 307LITE. [COLM 2024] LITE: Modeling Environmental Ecosystems with Multimodal Large Language Models
★ 14DiagrammerGPT. Official code repository for: DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning (COLM 2024)
★ 159explanation_based_rescaling. This repository hosts the dataset for explanation based rescaling
★ 3POPE. The official GitHub page for ''Evaluating Object Hallucination in Large Vision-Language Models''
★ 266SoM-LLaVA. [COLM-2024] List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
★ 146MMCoQA. Python
★ 31LLaMA-VID. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
★ 861Cheetah. Python
★ 354katna. Tool for automating common video key-frame extraction, video compression and Image Auto-crop/Image-resize tasks
★ 398MAVIS. [ICLR 2025] Mathematical Visual Instruction Tuning for Multi-modal Large Language Models
★ 156DenseFusion. DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
★ 159