This is your work, valued
SpatialEdit. [Official Repo] SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
★ 214UVCOM. [CVPR 2024] Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection
★ 117MambaTree. [NeurIPS2024 Spotlight] The official implementation of MambaTree: Tree Topology is All You Need in State Space Model
★ 112SemanticAC. [ICASSP 2023] SEMANTICAC: SEMANTICS-ASSISTED FRAMEWORK FOR AUDIO CLASSIFICATION
★ 5EasonXiao-888.
★ 1EasonXiao-888.github.io. HTML
★ 1RKDNet.
★ 1penguin-harness. 🐧 Automated Agent Builder on Your Desktop: Create Self-Evolving Agents in One Click
★ 221LiveEdit. [ECCV 2026] LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
★ 137Stable-Video-Infinity. [ICLR 26 Oral] Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
★ 2.5kBernini. Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
★ 1.2kLongLive-RAG. Official Implementation of LongLive-RAG: A general retrieval-augmented framework for long video generation.
★ 103drifting. Python
★ 485mammothmoda. Python
★ 333FastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kMeta-CoT. [CVPR 2026] Official code of the paper "Meta-CoT: Enhancing Granularity and Generalization in Image Editing"
★ 79DuQuant-v2. DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
★ 13Awesome-Efficient-dLLMs. 📚 A curated list of Awesome Efficient dLLMs Papers with Codes
★ 12QDLM. Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
★ 6HY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kOpenSpatial. Python
★ 94MotionLCM. [ ECCV 2024 ] MotionLCM: This repo is the official implementation of "MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model"
★ 462triattention. TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained GPUs.
★ 835SpatialEdit. [Official Repo] SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
★ 215JoyAI-Image. JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
★ 2.2kStreamDiffusionV2. StreamDiffusion, Live Stream APP
★ 533Helios. Helios: Real Real-Time Long Video Generation Model
★ 2kKiwi-Edit. A unified and fully open-source framework for instruction-guided and reference-guided video editing using natural language.
★ 312Utonia. [ICML'26] Official repository of Utonia: Toward One Encoder for All Point Clouds
★ 713Kimi-K2.5. Open Visual Agentic Intelligence
★ 2.3klingbot-world. Advancing Open-source World Models
★ 4.3kReasoning-Visual-World. Official repository for "Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", https://arxiv.org/abs/2601.19834
★ 100objaverse-xl. 🪐 Objaverse-XL is a Universe of 10M+ 3D Objects. Contains API Scripts for Downloading and Processing!
★ 1.3kOmni-Video. Python
★ 159sam-3d-objects. SAM 3D Objects
★ 7.2kReCamMaster. [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
★ 1.8kDSR_Suite. Jupyter Notebook
★ 74Wan-Move. [NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
★ 649VANS. [CVPR 2026] Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
★ 119VST. [ECCV2026] Visual Spatial Tuning
★ 201Emu3.5. Native Multimodal Models are World Learners
★ 1.5kMindOmni.
★ 3QeRL. [ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
★ 512LongLive. Long Video Gen Infrastructure
★ 2.5kHunyuanImage-3.0. HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
★ 3.2kMRVS_SOC. Python
★ 9VeOmni. VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
★ 2.1kFollowYourShape. [ICLR 2026] Follow-Your-Shape: This repo is the official implementation of "Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control"
★ 71Qwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kARC-Hunyuan-Video-7B. Structured Video Comprehension of Real-World Shorts
★ 240Long-RL. Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)
★ 727UniWorld. UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
★ 886RISEBench. [NIPS 2025 DB Oral] Official Repository of paper: Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
★ 155OmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kBagel. Open-source unified multimodal model
★ 6.1kTokLIP. TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
★ 236edge-tts. Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
★ 12kMindOmni. [NeurIPS2025] The official implementation of MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
★ 139Step1X-Edit. A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.
★ 2.2kVideoPainter. [SIGGRAPH2025] Official repo for paper "Any-length Video Inpainting and Editing with Plug-and-Play Context Control"
★ 626AnimeGamer. [ICCV 2025] AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
★ 346awesome-deep-reasoning. Collect every awesome work about r1!
★ 433x-flux. Python
★ 2.2kSana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6klmm-r1. Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.
★ 848HaploVLM. ICML2025
★ 63MM-EUREKA. MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
★ 771Megatron-LM. Ongoing research training transformer models at scale
★ 17kVoCo-LLaMA. [CVPR'2025] VoCo-LLaMA: This repo is the official implementation of "VoCo-LLaMA: Towards Vision Compression with Large Language Models".
★ 205VLM-R1. Solve Visual Understanding with Reinforced VLMs
★ 6kImage-Generation-CoT. [CVPR 2025] The First Investigation of CoT Reasoning (RL, TTS, Reflection) in Image Generation
★ 865RF-Solver-Edit. [🚀ICML 2025] "Taming Rectified Flow for Inversion and Editing" Using FLUX and HunyuanVideo for image and video editing!
★ 638lmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kOmniGen. OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
★ 4.3kECCV2022-RIFE. ECCV2022 - Real-Time Intermediate Flow Estimation for Video Frame Interpolation
★ 5.5kDynamicCity. [ICLR 2025 Spotlight] Official implementation for "DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes"
★ 2461d-tokenizer. This repo contains the code for 1D tokenizer and generator
★ 1.2kCogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kAutoregressive-Models-in-Vision-Survey. [TMLR 2025🔥] A survey for the autoregressive models in vision.
★ 805PhyGenBench. [ICML2025] The code and data of Paper: Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
★ 163mar. PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838
★ 1.9kConsistent123. [ACMMM 2024] Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
★ 25WritingAIPaper. Writing AI Conference Papers: A Handbook for Beginners
★ 3.9kAwesome-diffusion-model-for-image-processing. one summary of diffusion-based image processing, including restoration, enhancement, coding, quality assessment
★ 956HivisionIDPhotos. ⚡️HivisionIDPhotos: a lightweight and efficient AI ID photos tools. 一个轻量级的AI证件照制作算法。
★ 21knatural-instructions. Expanding natural instructions
★ 1kLumina-mGPT. Official Implementation of "Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining"
★ 647EditAnything. Edit anything in images powered by segment-anything, ControlNet, StableDiffusion, etc. (ACM MM)
★ 3.4kopen-instruct. AllenAI's post-training codebase
★ 3.8kCOVE. [NeurIPS 2024] COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
★ 26awesome-vlm-architectures. Famous Vision Language Models and Their Architectures
★ 1.3kSpeculativeDecodingPapers. 📰 Must-read papers and blogs on Speculative Decoding ⚡️
★ 1.3kAwesome-Efficient-LLM. A curated list for Efficient Large Language Models
★ 2kSEED-Voken. SEED-Voken: A Series of Powerful Visual Tokenizers
★ 1kMobileVLM. Strong and Open Vision Language Assistant for Mobile Devices
★ 1.4kchainlit. Build Conversational AI in minutes ⚡️
★ 12kEfficient-Multimodal-LLMs-Survey. Efficient Multimodal Large Language Models: A Survey
★ 387MambaTree. [NeurIPS2024 Spotlight] The official implementation of MambaTree: Tree Topology is All You Need in State Space Model
★ 112CoHD. The official implementation of A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
★ 27LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kLongLoRA. Code and documents of LongLoRA and LongAlpaca (ICLR 2024 Oral)
★ 2.7kLaVIT. LaVIT: Empower the Large Language Model to Understand and Generate Visual Content
★ 603awesome-diffusion-models-in-low-level-vision. A Repository for Diffusion-Model-related Papers in Low-level Vision
★ 556awesome-concealed-object-segmentation.
★ 347Video-LLaVA. 【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3.5kUniLSeg. [CVPR 2024] Official implementation of "Universal Segmentation at Arbitrary Granularity with Language Instruction"
★ 283VQA_to_multimodal_survey. Update 2020
★ 74FollowYourHandle. [WACV 2025] Follow-Your-Handle: This repo is the official implementation of "MagicStick: Controllable Video Editing via Control Handle Transformations"
★ 99UVCOM. [CVPR 2024] Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection
★ 117PKEF. [CIKM'23] Parallel Knowledge Enhancement based Framework for Multi-behavior Recommendation
★ 7EasonXiao-888.github.io. HTML
★ 1RKDNet.
★ 1NeurIPS2023_SOC. [NeurIPS 2023] The official implementation of SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation
★ 33iccv2023_RVOS_Challenge. [ICCV 2023 Workshop] The Official Implementation of The First Prize Solution for RVOS Competition
★ 14EasonXiao-888.
★ 1SemanticAC. [ICASSP 2023] SEMANTICAC: SEMANTICS-ASSISTED FRAMEWORK FOR AUDIO CLASSIFICATION
★ 5UniVTG. [ICCV 2023] UniVTG: Towards Unified Video-Language Temporal Grounding
★ 380Awesome-Cross-Modal-Video-Moment-Retrieval. 前沿论文持续更新--视频时刻定位 or 时域语言定位 or 视频片段检索。
★ 266CVPR2026-Papers-with-Code. CVPR 2026 论文和开源项目合集
★ 23kCVPR-2023-Papers.
★ 936Grounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kFollowYourPose. [AAAI 2024] Follow-Your-Pose: This repo is the official implementation of "Follow-Your-Pose : Pose-Guided Text-to-Video Generation using Pose-Free Videos"
★ 1.4kcs-self-learning. 计算机自学指南
★ 75kmmaction2. OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark
★ 5.1kCS-BAOYAN-2022. 计算机保研交流群(QQ群号:605176069)
★ 1.5kxduts. Xidian University TeX Suite 西安电子科技大学LaTeX套装
★ 1.1kBilibili-WenDao. Algorithm & DataSet for Machine Learning
★ 72