This is your work, valued
CameraCtrl. Python
★ 657EBLNet. Official code for ICCV2021 paper: Enhanced Boundary Learning for Glass-like Object Segmentation
★ 81Projects-CameraCtrl-II. HTML
★ 2projects-CameraCtrl. HTML
★ 1FastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9khermes-agent. The agent that grows with you
★ 222kdemoparser. Counter-Strike 2 replay parser for Python and JavaScript
★ 697LongLive-RAG. Official Implementation of LongLive-RAG: A general retrieval-augmented framework for long video generation.
★ 102tuna-2. Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
★ 739SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4kcli. The official Lark/Feishu CLI tool, maintained by the larksuite team — built for humans and AI Agents. Covers core business domains including Messenger, Docs, Base, Sheets, Calendar, Mail, Tasks, Meetings, and more, with 200+ commands and 20+ AI Agent Skills.
★ 16kcosmos-reason2. Cosmos-Reason2 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
★ 432ABot-PhysWorld. Python
★ 369JTA-Dataset. Python
★ 201InternUtopia. A simulation platform for versatile Embodied AI research and developments.
★ 1.3kUnrealClaude. Claude Code CLI integration for Unreal Engine 5.7 - Get AI coding assistance with built-in UE5.7 documentation context directly in the editor.
★ 867WorldScore. Official implementation for WorldScore: A Unified Evaluation Benchmark for World Generation
★ 302madpose. [CVPR 2025 Highlight] Official implementation of the solvers and estimators proposed in the paper "Relative Pose Estimation through Affine Corrections of Monocular Depth Priors"
★ 237VideoGPA. [ICML'26] VideoGPA is a self-supervised framework that enhances 3D consistency in Video Diffusion Models.
★ 70vipe. ViPE: Video Pose Engine for Geometric 3D Perception
★ 2.1klingbot-vla. A Pragmatic VLA Foundation Model
★ 1.7klingbot-depth. Masked Depth Modeling for Spatial Perception
★ 1.5kVLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3kMMSI-Bench. [ICLR 2026] MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
★ 106Structured3D. [ECCV'20] Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling
★ 678MagiAttention. A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
★ 893G2VLM. [CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
★ 347evo. Python package for the evaluation of odometry and SLAM
★ 4.3kDOVER. [ICCV 2023, Official Code] for paper "Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives". Official Weights and Demos provided.
★ 517sam-3d-objects. SAM 3D Objects
★ 7.2kDepth-Anything-3. Depth Anything 3
★ 6kMAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.8kVST. [ECCV2026] Visual Spatial Tuning
★ 200RollingForcing. [ICLR 2026] Official Repo for Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
★ 449LongLive. Long Video Gen Infrastructure
★ 2.5kOmniVinci. OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 675xfactor-nvs. Public code for XFactor: Introduces the first geometry-free model to achieve true self-supervised / pose-free Novel View Synthesis (NVS) by learning transferable latent camera pose representations.
★ 160Emu3.5. Native Multimodal Models are World Learners
★ 1.5kprolificdreamer. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation (NeurIPS 2023 Spotlight)
★ 1.6kDMD2. (NeurIPS 2024 Oral 🔥) Improved Distribution Matching Distillation for Fast Image Synthesis
★ 1.4kDepthLM_Official. [ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLM
★ 363Unify-Post-Training. Towards a Unified View of Large Language Model Post-Training
★ 211ddpo-pytorch. DDPO for finetuning diffusion models, implemented in PyTorch with LoRA support
★ 768SocialNavSUB. [CoRL 2025] VLM Benchmark for Social Navigation Scene Understanding
★ 24Awesome_Think_With_Images. Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
★ 1.5kWaver. Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.
★ 950YUME. The official code of Yume
★ 679VGGT-SLAM. VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
★ 1.1kNextStep-1. [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s Multimodal Intelligence team.
★ 693leapvo. [CVPR 2024] LEAP-VO: Long-term Effective Any Point Tracking for Visual Odometry
★ 254dinov3. Reference PyTorch implementation and models for DINOv3
★ 11kBLIP3o. Official implementation of BLIP3o-Series
★ 1.7kMoBA. MoBA: Mixture of Block Attention for Long-Context LLMs
★ 2.2kHunyuanWorld-1.0. Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
★ 2.9kSpaTrackerV2. [ICCV 2025] SpatialTrackerV2: 3D Point Tracking Made Easy
★ 985ao. PyTorch native quantization and sparsity for training and inference
★ 2.9kwaymo-open-dataset. Waymo Open Dataset
★ 3.4kCUT3R. Official implementation of Continuous 3D Perception Model with Persistent State
★ 1.5kEmbodiedOcc. [ICCV 2025] Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
★ 88taskonomy. Taskonomy: Disentangling Task Transfer Learning [Best Paper, CVPR2018]
★ 877omnidata. A Scalable Pipeline for Making Steerable Multi-Task Mid-Level Vision Datasets from 3D Scans [ICCV 2021]
★ 493radial-attention. [NeurIPS 2025] Radial Attention: O(nlogn) Sparse Attention with Energy Decay for Long Video Generation
★ 605OpenSfM. Open source Structure-from-Motion pipeline
★ 3.8kd2-net. D2-Net: A Trainable CNN for Joint Description and Detection of Local Features
★ 847dvgformer. Code for our paper: Learning Camera Movement Control from Real-World Drone Videos
★ 36FlashDepth. The official implementation of ICCV'25 paper "FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution"
★ 395sekai-codebase. [NeurIPS 2025] Sekai: A Video Dataset towards World Exploration
★ 302Tar. [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
★ 202openpi. Python
★ 13kOmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kMoGe. [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
★ 2.7kuco3d. Uncommon Objects in 3D dataset
★ 1.3kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kkubric. A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
★ 2.8kSeedVR. Repo for SeedVR2 (ICLR2026) & SeedVR (CVPR2025 Highlight)
★ 1.3kWorldMem. [NeurIPS 2025] WorldMem: Long-term Consistent World Simulation with Memory
★ 381LAPA. [ICLR 2025] LAPA: Latent Action Pretraining from Videos
★ 562UniGeo. UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
★ 136RoboMaster. [ICLR’26] Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
★ 107OpenUni. Python
★ 189FlowMo. Official PyTorch implementation of FlowMo.
★ 117LoopNav. Python
★ 16reverse-engineering-gemma-3n. Reverse Engineering Gemma 3n: Google's New Edge-Optimized Language Model
★ 279co-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5kMMaDA. MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
★ 1.7kBagel. Open-source unified multimodal model
★ 6.1kFoundationPose. [CVPR 2024 Highlight] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
★ 3.5kflow_grpo. [NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL
★ 2.4kRealCam-I2V. Python
★ 60CamI2V. official repo of paper for "CamI2V: Camera-Controlled Image-to-Video Diffusion Model"
★ 171DiffSynth-Studio. Enjoy the magic of Diffusion models!
★ 13kVoyager. An Open-Ended Embodied Agent with Large Language Models
★ 7.1kMagma. [CVPR 2025] Magma: A Foundation Model for Multimodal AI Agents
★ 1.9kUI-TARS. Pioneering Automated GUI Interaction with Native Agents
★ 11kJanus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kvila-u. [ICLR 2025] VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
★ 425T2I-R1. [NeurIPS 2025] T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
★ 433ml-tarflow. Python
★ 341CameraBench. [NeurIPS 2025 Spotlight] Towards Understanding Camera Motions in Any Video
★ 305awesome-alignment-of-diffusion-models. [ACM Computing Surveys] The collection of awesome papers on alignment of diffusion models.
★ 430REPA-E. [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers
★ 511MetaSpatial. [ICLR 2026] MetaSpatial leverages reinforcement learning to enhance 3D spatial reasoning in vision-language models (VLMs), enabling more structured, realistic, and adaptive scene generation for applications in the metaverse, AR/VR, and game development.
★ 322HunyuanVideo-I2V. HunyuanVideo-I2V: A Customizable Image-to-Video Model based on HunyuanVideo
★ 1.8kFramePack. Lets make video diffusion practical!
★ 17kperception_models. State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
★ 2.3kSeedVR. [CVPR2025 Highlight] SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
★ 106ml-matrix3d. [CVPR 2025 Highlight] Matrix3D: Large Photogrammetry Model All-in-One
★ 597Geo4D. [ICCV 2025 Highlight] Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
★ 437metamorph. Code for MetaMorph Multimodal Understanding and Generation via Instruction Tuning
★ 235ARD. [CVPR 2025 Oral] PyTorch re-implementation for Autoregressive Distillation of Diffusion Transformers (ARD).
★ 144SuperPoint. Efficient neural feature detector and descriptor
★ 2.5kGigaTok. [ICCV 2025] Official repo for "GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation"
★ 204finetrainers. Scalable and memory-optimized training of diffusion models
★ 1.4kDDT. [CVPR 2026] DDT: Decoupled Diffusion Transformer
★ 407LLaMA-Mesh. Unifying 3D Mesh Generation with Language Models
★ 1.2kPixelFlow. Pixel-Space Generative Models
★ 316deep-rl-class. This repo contains the Hugging Face Deep Reinforcement Learning Course.
★ 5kunimatch. [TPAMI'23] Unifying Flow, Stereo and Depth Estimation
★ 1.4kmuggled_dpt. Muggled DPT: Depth estimation without the magic
★ 118croco. Python
★ 508PAR. [CVPR2025 Highlight] PAR: Parallelized Autoregressive Visual Generation. https://yuqingwang1029.github.io/PAR-project
★ 186GoT. Official repository of "GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing"
★ 317AnimeGamer. [ICCV 2025] AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
★ 346anycam. Official repository for "AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos" (CVPR 2025)
★ 303DAPO. An Open-source RL System from ByteDance Seed and Tsinghua AIR
★ 1.8kSegAnyMo. [CVPR 2025] Code for Segment Any Motion in Videos
★ 485dso. [ICCV 2025] DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness
★ 189DeepSeek-Math. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
★ 3.4kFAR. Code for: "Long-Context Autoregressive Video Modeling with Next-Frame Prediction"
★ 311Qwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
★ 4.1kOpenRLHF. An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
★ 9.9kcarve3d. Code for Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning
★ 37fractalgen. PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
★ 1.2kRDLM. Official Code Repository for the paper "Continuous Diffusion Model for Language Modeling" (NeurIPS 2025).
★ 74UI-TARS-desktop. The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
★ 38kbd3lms. [ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
★ 1kLLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kFLARE. Python
★ 721Video-T1. [ICCV 2025] Video-T1: Test-Time Scaling for Video Generation
★ 317cosmos-reason1. Cosmos-Reason1 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
★ 952stable-virtual-camera. Stable Virtual Camera: Generative View Synthesis with Diffusion Models
★ 1.6kReCamMaster. [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
★ 1.8kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kStep-Video-T2V. Python
★ 3.2kVideo-Depth-Anything. [CVPR 2025 Highlight] Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
★ 2kUniTok. [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding
★ 529EasyR1. EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kscaling-with-vocab. [NeurIPS-2024] 📈 Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies https://arxiv.org/abs/2407.13623
★ 112nabla-gfn. Official Implementation of Nabla-GFlowNet (ICLR 2025)
★ 28open-infra-index. Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
★ 8kwonderland. Python
★ 166WonderJourney. Python
★ 770DiffusionAsShader. [SIGGRAPH 2025] Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
★ 824SkyReels-V1. SkyReels V1: The first and most advanced open-source human-centric video foundation model
★ 2.7kGS-DiT. Source code for paper GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking
★ 55WonderWorld. Code release for https://kovenyu.com/WonderWorld/
★ 739SynCamMaster. [ICLR'25] SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
★ 693diamond. DIAMOND (DIffusion As a Model Of eNvironment Dreams) is a reinforcement learning agent trained in a diffusion world model. NeurIPS 2024 Spotlight.
★ 2.1kac3d. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
★ 164Lumina-Video. Python
★ 416VideoAlign. [NeurIPS 2025] Improving Video Generation with Human Feedback
★ 489DB. A PyTorch implementation of "Real-time Scene Text Detection with Differentiable Binarization".
★ 2.3ktokenizers. 💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
★ 11kopen-r1. Fully open reproduction of DeepSeek-R1
★ 26k1Prompt1Story. 🔥ICLR 2025 (Spotlight) One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
★ 320TokenFlow. [CVPR 2025] 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".
★ 464Vchitect-2.0. Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
★ 922Min-SNR-Diffusion-Training. [ICCV 2023] Efficient Diffusion Training via Min-SNR Weighting Strategy
★ 269GameFactory. [ICCV 2025] GameFactory: Creating New Games with Generative Interactive Videos
★ 495SONAR. SONAR, a new multilingual and multimodal fixed-size sentence embedding space, with a full suite of speech and text encoders and decoders.
★ 900LlamaV-o1. [ACL 2025 🔥] Rethinking Step-by-step Visual Reasoning in LLMs
★ 307large_concept_model. Large Concept Models: Language modeling in a sentence representation space
★ 2.4kMegaSynth. Code for MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data (CVPR 2025)
★ 205glomap. [DEPRECATED] GLOMAP - Global Structured-from-Motion Revisited
★ 2.4kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11k1xgpt. world modeling challenge for humanoid robots
★ 564Sa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.6kVisionReward. [AAAI 2026] VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
★ 422REPA. [ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
★ 1.7kLightningDiT. [CVPR 2025 Oral] Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
★ 1.5kcausalfusion. Python
★ 197sparse_coding. Using sparse coding to find distributed representations used by neural networks.
★ 307ml-depth-pro. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.
★ 5.6kUniDepth. Universal Monocular Metric Depth Estimation
★ 1.2kZoeDepth. Metric depth estimation from a single image
★ 2.8kMetric3D. The repo for "Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image" and "Metric3Dv2: A Versatile Monocular Geometric Foundation Model..."
★ 2.3kDeepSeek-V3. Python
★ 104kPromptDA. [CVPR 2025] Prompt Depth Anything
★ 1.1kVideoVAEPlus. [ICCV 2025] VideoVAE+: Large Motion Video Autoencoding with Cross-modal Video VAE
★ 410open-muse. Open reproduction of MUSE for fast text2image generation.
★ 358DROID-Splat. End-to-End SLAM with camera calibration, monocular prior integration and dense Rendering
★ 416VideoDPO. Official Implementation of VideoDPO
★ 169VidTok. a family of versatile and state-of-the-art video tokenizers.
★ 454NOVA. [ICLR 2025] Autoregressive Video Generation without Vector Quantization
★ 657MVAR. Python
★ 71HumanVid. [NeurIPS D&B Track 2024] Official implementation of HumanVid
★ 349DeepSeek-VL2. DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
★ 5.3kStreamDiffusion. StreamDiffusion: A Pipeline-Level Solution for Real-Time Interactive Generation
★ 11kgenex. Generative World Explorer
★ 167Owl.
★ 52streamv2v. Official Pytorch implementation of StreamV2V.
★ 546distillnerf. [NeurIPS 2024] DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features
★ 38flow_matching. A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
★ 4.7k