This is your work, valued
Open-LLaVA-NeXT. An open-source implementation for training LLaVA-NeXT.
★ 439DALN. [CVPR2022] Official implementation of DALN.
★ 93DDB. [NeurIPS 2022 Spotlight] Official implement of Deliberated Domain Bridging for Domain Adaptive Semantic Segmentation
★ 70VLMEvalKit. Python
★ 2SCOPE. Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
★ 32SkillFlow. Python
★ 41OneResearchClaw. Any research. One Claw. 🦞 From any materials to research with fully autonomous & skill-driven researcher.
★ 440Video-MME-v2. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
★ 369WildClawBench. An in-the-wild benchmark for AI agents in the OpenClaw Environment.
★ 498MoDA. An hardware-aware Efficient Implementation for "Mixture-of-Depths Attention".
★ 274Vision-DeepResearch. [ICML 2026] Multimodal deep-research MLLM and benchmark. The first long-horizon multimodal deep-research MLLM, extending the number of reasoning turns to dozens and the number of search-engine interactions to hundreds.
★ 659youtu-vl. Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
★ 169manim. Animation engine for explanatory math videos
★ 89kVeOmni. VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
★ 2.1kAGILE. [ICLR2026] Official Implement of "Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models"
★ 117DAEDAL. [ICLR 2026] Official repository of "Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models"
★ 173ScaleCap. (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’
★ 60Awesome-Interleaving-Reasoning. Interleaving Reasoning: Next-Generation Reasoning Systems for AGI
★ 280MM-EUREKA. MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
★ 771SeekWorld. The first attempt to replicate o3-like visual clue-tracking reasoning capabilities.
★ 64Seed1.5-VL. Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
★ 1.6kVCR-Bench. VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
★ 37Kimi-VL. Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities
★ 1.2kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kDAPO. An Open-source RL System from ByteDance Seed and Tsinghua AIR
★ 1.8kUniToken. [CVPRW 2025] UniToken is an auto-regressive generation model that combines discrete and continuous representations to process visual inputs, making it easy to integrate both visual understanding and image generation tasks seamlessly.
★ 106Vision-R1. [ICLR2026] This is the first paper to explore how to effectively use R1-like RL for MLLMs and introduce Vision-R1, a reasoning MLLM that leverages cold-start initialization and RL training to incentivize reasoning capability.
★ 1.6kLIMO. [COLM 2025] LIMO: Less is More for Reasoning
★ 1.1kVisualThinker-R1-Zero. Explore the Multimodal “Aha Moment” on 2B Model
★ 624cognitive-behaviors. Python
★ 224MM-Eureka-V0. MM-Eureka V0 also called R1-Multimodal-Journey, Latest version is in MM-Eureka
★ 325EasyR1. EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1kVisual-RFT. Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’
★ 2.3kVisualPerceptionToken. Python
★ 136Awesome-RL-based-Reasoning-MLLMs. This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
★ 1.4kR1-V. Witness the aha moment of VLM with less than $3.
★ 4.1kOpen-R1-Video. ✨First Open-Source R1-like Video-LLM [2025/02/18]
★ 382Light-A-Video. [ICCV 2025] Light-A-Video: Training-free Video Relighting via Progressive Light Fusion
★ 517VideoRoPE. [ICML 2025 Oral] An official implementation of VideoRoPE & VideoRoPE++
★ 223cosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kTransNetV2. TransNet V2: Shot Boundary Detection Neural Network
★ 1kMulberry. [NIPS'25 Spotlight] Mulberry, an o1-like Reasoning and Reflection MLLM Implemented via Collective MCTS
★ 1.2kVisualSketchpad. Codes for Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
★ 287Awesome-LLM-Reasoning. From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
★ 3.7kiTerm2-Color-Schemes. Over 450 terminal color schemes/themes for iTerm/iTerm2. Includes ports to Terminal, Konsole, PuTTY, Xresources, XRDB, Remmina, Termite, XFCE, Tilda, FreeBSD VT, Terminator, Kitty, MobaXterm, LXTerminal, Microsoft's Windows Terminal, Visual Studio, Alacritty, Ghostty, and many more
★ 27kms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kMinerU. Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
★ 76kSAM2Long. [ICCV 2025] SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
★ 568LLaVA-CoT. [ICCV 2025] LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning
★ 2.1kleetcode-master. 《代码随想录》LeetCode 刷题攻略:200道经典题目刷题顺序,共60w字的详细图解,视频难点剖析,50余张思维导图,支持C++,Java,Python,Go,JavaScript等多语言版本,从此算法学习不再迷茫!🔥🔥 来看看,你会发现相见恨晚!🚀
★ 62kUGround. [ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agents
★ 317Modality-Integration-Rate. [ICCV 2025] The official code of the paper "Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate".
★ 113Evaluation-Multimodal-LLMs-Survey. A Survey on Benchmarks of Multimodal Large Language Models
★ 157WindowsAgentArena. Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
★ 883UFO. UFO³: Weaving the Digital Agent Galaxy
★ 9.4kSeeAct. [ICML'24] SeeAct is a system for generalist web agents that autonomously carry out tasks on any given website, with a focus on large multimodal models (LMMs) such as GPT-4V(ision).
★ 850Qwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kVITA. ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
★ 2.5kGrounded-SAM-2. Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2
★ 3.7ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kMindSearch. 🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT)
★ 6.9kSEED-Voken. SEED-Voken: A Series of Powerful Visual Tokenizers
★ 1k1d-tokenizer. This repo contains the code for 1D tokenizer and generator
★ 1.2klmms-finetune. A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave, llava-next-video, llava-onevision, llama-3.2-vision, qwen-vl, qwen2-vl, phi3-v etc.
★ 373Awesome-Dataset-Distillation. A curated list of awesome papers on dataset distillation and related applications.
★ 2kSPPO. The official implementation of Self-Play Preference Optimization (SPPO)
★ 590STIC. Enhancing Large Vision Language Models with Self-Training on Image Comprehension.
★ 68conv-llava. Python
★ 128DocGenome. DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Models
★ 156cambrian. Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2kMLVU. 🔥🔥MLVU: Multi-task Long Video Understanding Benchmark
★ 266Prism. A Framework for Decoupling and Assessing the Capabilities of VLMs
★ 44Uni-MoE. Uni-MoE: Lychee's Large Multimodal Model Family.
★ 1.1kVILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8kSTAR. STAR: Scale-wise Text-to-image generation via Auto-Regressive representations
★ 150VideoLLaMA2. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
★ 1.3kdeprecated-generative-ai-python. This SDK is now deprecated, use the new unified Google GenAI SDK.
★ 2.3klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kLongVideoBench. [Neurips 24' D&B] Official Dataloader and Evaluation Scripts for LongVideoBench.
★ 134Video-MME. ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
★ 788MotionClone. [ICLR 2025] Official implementation of MotionClone: Training-Free Motion Cloning for Controllable Video Generation
★ 516ShareGPT4Omni. ShareGPT4Omni: Towards Building Omni Large Multi-modal Models with Comprehensive Multi-modal Annotations
★ 10ShareGPT4Video. [NeurIPS 2024] An official implementation of "ShareGPT4Video: Improving Video Understanding and Generation with Better Captions"
★ 1.1kShareGPT4V. [ECCV 2024] ShareGPT4V: Improving Large Multi-modal Models with Better Captions
★ 259MovieChat. [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
★ 706MiniCPM-V. A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
★ 26kMGM. Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
★ 3.3kYi-1.5. Yi-1.5 is an upgraded version of Yi, delivering stronger performance in coding, math, reasoning, and instruction-following capability.
★ 559Open-LLaVA-NeXT. An open-source implementation for training LLaVA-NeXT.
★ 439Video-Bench. A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models!
★ 140MambaOut. MambaOut: Do We Really Need Mamba for Vision? (CVPR 2025)
★ 2.7kLLaVA-NeXT. Python
★ 4.7kCVRR-Evaluation-Suite. [CVPRW-25 MMFM] Official repository of paper titled "How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs".
★ 50LLaVA-pp. 🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
★ 841lmdeploy. LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
★ 8kLaVi-Bridge. [ECCV 2024] Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
★ 300ELLA. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
★ 1.3kSoraReview. The official GitHub page for the review paper "Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models".
★ 504Portrait-Mode-Video. Video dataset dedicated to portrait-mode video recognition.
★ 57Open-Sora-Dataset. Python
★ 115GenPercept. [ICLR2025] GenPercept: Diffusion Models Trained with Large Data Are Transferable Visual Models
★ 229Open-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kMathVerse. [ECCV 2024] Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
★ 183MMStar. [NeurIPS 2024] This repo contains evaluation code for the paper "Are We on the Right Way for Evaluating Large Vision-Language Models"
★ 215Agent-FLAN. [ACL2024 Findings] Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
★ 361vida. [ICLR 2024] ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
★ 78MobileVLM. Strong and Open Vision Language Assistant for Mobile Devices
★ 1.4kOpen-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kTinyLLaVA_Factory. A Framework of Small-scale Large Multimodal Models
★ 995InternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kpromptbench. A unified evaluation framework for large language models
★ 2.8kprismatic-vlms. A flexible and efficient codebase for training visually-conditioned language models (VLMs)
★ 1kMouSi.
★ 75MoE-LLaVA. 【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models
★ 2.3kYi. A series of large language models trained from scratch by developers @01-ai
★ 7.8kChinese-CLIP. Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6kllama-moe. ⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024)
★ 1kmistral-inference. Official inference library for Mistral models
★ 11kmixtral_spliter. Converting Mixtral-8x7B to Mixtral-[1~7]x7B
★ 22Aurora. The official codes for "Aurora: Activating chinese chat capability for Mixtral-8x7B sparse Mixture-of-Experts through Instruction-Tuning"
★ 261Rein. [CVPR 2024] Official implement of <Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation>
★ 411LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kVLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3kxtuner. A Next-Generation Training Engine Built for Ultra-Large MoE Models
★ 5.2kT-Eval. [ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step
★ 312transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kEgoThink. [CVPR'24 Highlight] The official code and data for paper "EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models"
★ 65trl. Train transformer language models with reinforcement learning.
★ 19kSEED-Bench. (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
★ 366honeybee. Official implementation of project Honeybee (CVPR 2024)
★ 469DCI. Densely Captioned Images (DCI) dataset repository.
★ 197MixtralKit. A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI
★ 770LLMGA. This project is the official implementation of 'LLMGA: Multimodal Large Language Model based Generation Assistant', ECCV2024 Oral
★ 395Prompt-Highlighter. [CVPR 2024] Prompt Highlighter: Interactive Control for Multi-Modal LLMs
★ 159Track-Anything. Track-Anything is a flexible and interactive tool for video object tracking and segmentation, based on Segment Anything, XMem, and E2FGVI.
★ 7kCogVLM. a state-of-the-art-level open visual language model | 多模态预训练模型
★ 6.7kMonkey. Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)
★ 2kLLaVA-Grounding. Python
★ 405LLaMA-VID. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
★ 861Chat-UniVi. [CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
★ 943LanguageBind. 【ICLR 2024🔥】 Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
★ 883Video-LLaVA. 【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3.5kmaestro. streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
★ 2.7ksegment-caption-anything. [CVPR'24] The repository provides code for running inference and training for "Segment and Caption Anything" (SCA) , links for downloading the trained model checkpoints, and example notebooks / gradio demo that show how to use the model.
★ 233RLHF-V. [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
★ 310ViP-LLaVA. [CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
★ 339MMMU. This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
★ 590ShareGPT4V-colab. Jupyter Notebook
★ 30PixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kBaichuan2. A series of large language models developed by Baichuan Intelligent Technology
★ 4.1kGPT4-V-selenium. Python
★ 4V3Det. Python
★ 121LLaVA-Interactive-Demo. LLaVA-Interactive-Demo
★ 380Woodpecker. ✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 649DETRDistill. [ICCV2023] DETRDistill: A Universal Knowledge Distillation Framework for DETR-families
★ 67DiT. Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
★ 8.7kSoM. [arXiv 2023] Set-of-Mark Prompting for GPT-4V and LMMs
★ 1.6kLLaMA2-Accessory. An Open-source Toolkit for LLM Development
★ 2.8ktree-sitter-python. Python grammar for tree-sitter
★ 560starcoder. Home of StarCoder: fine-tuning & inference!
★ 7.5kLISA. Project Page for "LISA: Reasoning Segmentation via Large Language Model"
★ 2.7kInternLM-XComposer. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2.9kLaVIN. [NeurIPS 2023] Official implementations of "Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models"
★ 523SolidUI. one sentence generates any graph
★ 658nougat. Implementation of Nougat Neural Optical Understanding for Academic Documents
★ 10kcodellama. Inference code for CodeLlama models
★ 16kOpenLLM. Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
★ 12kshikra. Python
★ 814Qwen-VL. The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
★ 6.7kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kstanford_alpaca. Code and documentation to train Stanford's Alpaca models, and generate the data.
★ 30kllama. Inference code for Llama models
★ 60kOtter. 🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
★ 3.4kopencompass. OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
★ 7.2kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26klagent. A lightweight framework for building LLM-based agents
★ 2.3kInternLM. Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
★ 7.3kCoDeF. [CVPR'24 Highlight] Official PyTorch implementation of CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
★ 4.8kStableVideo. [ICCV 2023] StableVideo: Text-driven Consistency-aware Diffusion Video Editing
★ 1.4kaccelerate. 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
★ 9.8kSubject-Diffusion. Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
★ 317svdiff-pytorch. Implementation of "SVDiff: Compact Parameter Space for Diffusion Fine-Tuning"
★ 386TokenFlow. Official Pytorch Implementation for "TokenFlow: Consistent Diffusion Features for Consistent Video Editing" presenting "TokenFlow" (ICLR 2024)
★ 1.7kprompt-to-prompt. Jupyter Notebook
★ 3.5ke4t-diffusion. Implementation of Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models
★ 324oft. Official implementation of "Controlling Text-to-Image Diffusion by Orthogonal Finetuning".
★ 300Semantic-SAM. [ECCV 2024] Official implementation of the paper "Semantic-SAM: Segment and Recognize Anything at Any Granularity"
★ 2.9kDTP. [ICCV 2023] Official implement of <Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement>
★ 71LAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kfastcomposer. [IJCV] FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
★ 715Awesome-Controllable-T2I-Diffusion-Models. A collection of resources on controllable generation with text-to-image diffusion models.
★ 1.1kICCV2025-Papers-with-Code. ICCV 2025 论文和开源项目合集
★ 2.9kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kDreambooth-Stable-Diffusion. Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
★ 7.7kmmagic. OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
★ 7.4kPersonalize-SAM. Personalize Segment Anything Model (SAM) with 1 shot in 10 seconds
★ 1.7kFreeDrag. [CVPR 2024] Official implementation of FreeDrag: Feature Dragging for Reliable Point-based Image Editing
★ 418DragDiffusion. [CVPR2024, Highlight] Official code for DragDiffusion
★ 1.3kDragonDiffusion. ICLR 2024 (Spotlight)
★ 788DragDiffusion. Implementation of DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing
★ 224Awesome-DragGAN. Awesome-DragGAN: A curated list of papers, tutorials, repositories related to DragGAN
★ 83