This is your work, valued
DeepLiDAR. Deep Surface Normal Guided Depth Prediction for Outdoor Scene from Sparse LiDAR Data and Single Color Image (CVPR 2019)
★ 260NeuS-HSR. Looking Through the Glass: Neural Surface Reconstruction Against High Specular Reflections (CVPR 2023)
★ 63SlimConv. Reducing Channel Redundancy in Convolutional Neural Networks by Features Recombining (TIP 2021)
★ 20NeuSmoke. NeuSmoke: Efficient Smoke Reconstruction and View Synthesis with Neural Transportation Fields (SIGGRAPH Asia 2024)
★ 19GauSim. A real-world autonomous driving simulator based on 3D Gaussian Splatting for scene augmentation
★ 16NeRC. NeRC: Rendering Planar Caustics by Learning Implicit Neural Representations (TVCG 2023)
★ 10Denoise. Fast video denoising
★ 7RCNeRF. Rendering Real-World Unbounded Scenes with Cars by Learning Positional Bias (TVC 2023)
★ 4RDNeRF. RDNeRF: Relative Depth Guided NeRF for Dense Free View Synthesis (TVC 2023)
★ 4slam-system. A simple SLAM system
★ 3Voxurf. [ ICLR 2023 Spotlight ] Pytorch implementation for "Voxurf: Voxel-based Efficient and Accurate Neural Surface Reconstruction"
★ 2DeepBlindness. Fast Blindness Map Estimation and Blindness Type Classification for Outdoor Scene from Single Color Image
★ 2surface-normal. This is a tool.code for CVPR2019 paper 1899: Deep Surface Normal Guided Depth Prediction for Outdoor Secene from Sparce Lidar Data and Single Color Image.
★ 1sam3.cpp. Fast state-of-the-art image and video segmentation in portable C/C++
★ 340lingbot-depth. Masked Depth Modeling for Spatial Perception
★ 1.5kqubit. Dual-arm desk robot which can help automate various simple tasks on any desk
★ 100GenRecon. Python
★ 910reloc3r. [CVPR 2025] Relative camera pose estimation and visual localization with Reloc3r
★ 321terrain-diffusion. Procedural generation with diffusion models (SIGGRAPH '26)
★ 1.3kblender-mcp. Open-source MCP to use Blender with any LLM
★ 25kEDGS. [CVPR 2026] A PyTorch implementation of the paper "EDGS: Eliminating Densification for Efficient Convergence of 3DGS"
★ 727Spatial-TTT. [ECCV 2026] Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
★ 243depth-anything-tensorrt. C++ TensorRT implementation of Depth-Anything V1, V2
★ 482minimind-o. 🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!
★ 2.2kgluemap. GLUEMAP: Global Structure-from-Motion Meets Feedforward Reconstruction
★ 317salad. Optimal Transport Aggregation for Visual Place Recognition
★ 369UniPR-3D. [ECCV 2026] UniPR-3D
★ 33mesh_align. A tool for rigid mesh alignment written in Python
★ 6GHOST. Python
★ 14RoboVerse. RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
★ 1.8kUniCorrn. [CVPR 2026] UniCorrn: Unified Correspondence Transformer Across 2D and 3D
★ 215unitree_mujoco. C++
★ 1.1kLongStream. Python
★ 91Pi-Long. Code implementation of Pi-Long
★ 193lingbot-map. A feed-forward 3D foundation model for reconstructing scenes from streaming data
★ 16kFlashVGGT. Accelerate VGGT with efficient desciptor-based global attention
★ 96SAM3-TENSORRT-PYTHON. Python
★ 29HY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kwanderland. [CVPR 2026 Highlight] Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
★ 60FlashSAC. FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
★ 422NavigationScaling. Official implementation of the RA-L 2026 paper "Data Scaling for Navigation in Unknown Environments"
★ 2GPS_IMU_Kalman_Filter. Fusing GPS, IMU and Encoder sensors for accurate state estimation.
★ 6503D-Fixer. [CVPR 2026] 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
★ 96sub2api. Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
★ 35kclaude-relay-service. CRS-自建Claude Code镜像,一站式开源中转服务,让 Claude、OpenAI、Gemini、Droid 订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
★ 12kvggt-onnx. ONNX models of VGGT
★ 71video_to_world. Our method reconstructs 3D worlds from video diffusion models using non-rigid alignment to resolve inherent 3D inconsistencies in the generated sequences.
★ 280edict. 🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
★ 16kQwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kawesome-openclaw-skills. The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
★ 52kZipMap. [CVPR 2026] ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
★ 483Evo-RL. We release Evo-RL, the opensource real-world offline RL on So-101 and AgileX PiPER for easier reproduction.
★ 725EDULITE_A3. Lightweight fully open-source 6-DOF robotic arm
★ 154tttLRM. [CVPR 2026 Highlight] tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction
★ 449HDMI. Python
★ 630claude-code-local-for-vscode. Claude Code VS Code extension patched for Force Local mode — run CLI locally, proxy file ops to remote server via VS Code Remote SSH
★ 44mimiclaw. MimiClaw: Harness on a $5 chip. No OS(Linux). No Node.js. No Mac mini. No Raspberry Pi. No VPS. Hardware agents OS.
★ 5.6kSAM3-TensorRT. TensorRT Engine for SAM-3 model by Meta AI
★ 190vllm-omni. A framework for efficient model inference with omni-modality models
★ 5.7kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 1InfiniteVGGT. The official implementation of InfiniteVGGT
★ 379NavDP. [ICRA 2026] NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
★ 7193D-TransUNet. This is the official repository for the paper "3D TransUNet: Advancing Medical Image Segmentation through Vision Transformers"
★ 315supervoxel-loss. Efficient connectivity-preserving loss function for training neural networks to perform instance segmentation.
★ 13IGGT_official. [ICLR'26] IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
★ 427humanoid-dexart-retarget. Retargeting of whole-body human motion to humanoid robots for dexterous manipulation of articulated objects.
★ 34video2robot. End-to-end pipeline converting generative videos (Veo, Sora) to humanoid robot motions
★ 716DynamicVerse. [NeurIPS 2025]"DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling"
★ 101robotic_world_model. Repository for our papers: Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics and Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots
★ 663newton. An open-source, GPU-accelerated physics simulation engine built upon NVIDIA Warp, specifically targeting roboticists and simulation researchers.
★ 5.3kInternVLA-M1. InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
★ 418sam-3d-body. The repository provides code for running inference with the SAM 3D Body Model (3DB), links for downloading the trained model checkpoints and datasets, and example notebooks that show how to use the model.
★ 3.4kpico. [CVPR 2025] PICO: Reconstructing 3D People In Contact with Objects
★ 70Depth-Anything-3. Depth Anything 3
★ 6khoibench111. Python
★ 6GMR. [arXiv 2025] GMR: General Motion Retargeting. Retarget human motions into diverse humanoid robots in real time on CPU. Retargeter for TWIST.
★ 72egoallo. Estimating Body and Hand Motion in an Ego-sensed World
★ 289TWIST. [CoRL 2025] TWIST: Teleoperated Whole-Body Imitation System
★ 794GMR. [ICRA 2026] GMR: General Motion Retargeting. Retarget human motions into diverse humanoid robots in real time on CPU. Retargeter for TWIST.
★ 2.5kAwesomeWorldModels. A Comprehensive Survey on World Models for Embodied AI
★ 339Awesome-World-Models. A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
★ 3.2kEgo-VCP. Ego-Vision World Model for Humanoid Contact Planning
★ 188lerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kbstro. The official code for BSTRO in paper: Capturing and Inferring Dense Full-Body Human-Scene Contact, CVPR2022
★ 103composition_rendering. Python
★ 109SynCamMaster. [ICLR'25] SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
★ 693Human3R. [ICLR 2026] An unified model for 4D human-scene reconstruction
★ 519TTT3R. [ICLR 2026] A simple state update rule to enhance length generalization for CUT3R
★ 711smplreg. Registration between a reconstructed point cloud and an estimated SMPL mesh.
★ 333dgs-avatar-release. 3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting
★ 436lyra. Project Lyra: Open Generative 3D World Models
★ 2.2kQwen3-Omni. Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
★ 3.9kGeoSVR. [NeurIPS'25 Spotlight] GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
★ 193map-anything. MapAnything: Universal Feed-Forward Metric 3D Reconstruction
★ 3.6kVideoMimic. Visual Imitation Enables Contextual Humanoid Control. CoRL 2025, Best Student Paper Award.
★ 819OmniWorld. [ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
★ 486CUT3R. Official implementation of Continuous 3D Perception Model with Persistent State
★ 1.5kSpatialVID. [CVPR 2026] SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
★ 589sail-recon. [3DV 2026 Oral] Official Repo of "SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization"
★ 299DreamLifting. The code implementation for the paper "DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation".
★ 30Worldeval. Python
★ 29Awesome-World-Models. A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related websites.
★ 1.9krsrd. Code for "Robot See Robot Do" presented at CoRL 2024!
★ 163MV-Adapter. [ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
★ 1.3ktextured_gaussians. Code release for CVPR 2025 paper "Textured Gaussians for Enhanced 3D Scene Appearance Modeling".
★ 106Orbit. Unified framework for robot learning built on NVIDIA Isaac Sim. [ALR custom version]
★ 8LongSplat. [ICCV 2025] LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
★ 796mvtracker. [ICCV 2025 Oral] MVTracker: Multi-view 3D Point Tracking
★ 512SpatialGen. [3DV 2026] SpatialGen: Layout-guided 3D Indoor Scene Generation
★ 403FastMesh. [3DV 2026] FastMesh: Efficient Artistic Mesh Generation via Component Decoupling
★ 138Embodied-R1. Official code for "Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation" (ICLR2026)
★ 152pyroki. A Modular Toolkit for Robot Kinematic Optimization
★ 1.6kRynnEC. RynnEC: Bringing MLLMs into Embodied World
★ 391x-humanoid-robomind.github.io. JavaScript
★ 132PanoSplatt3R. [ICCV 2025] PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
★ 33real2render2real. [CoRL 2025] Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
★ 367dsrl. Official implementation for DSRL, Steering Your Diffusion Policy with Latent Space Reinforcement Learning (CoRL 2025)
★ 208Awesome-Flow-RL-Papers. A collection of paper/projects that trains flow matching model/policies via RL.
★ 408ReinFlow. [NeurIPS 2025] Flow x RL. "ReinFlow: Fine-tuning Flow Policy with Online Reinforcement Learning". Support VLAs e.g., Pi0, Pi0.5, GR00TN1.5. Fully open-sourced.
★ 349GLOVER. This is the official code repo for GLOVER and GLOVER++.
★ 58OSAD_Net. Pytorch implementation of One-Shot Affordance Detection
★ 654DNeX. 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
★ 840scenefun3d. SceneFun3D ToolKit
★ 181unsup-affordance. Python
★ 152splatt3r. Official repository for Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
★ 797RoboSplat. [RSS 2025] Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation
★ 197Awesome-3DGS-Applications. 【TPAMI 2026】A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
★ 393Vid2Sim. [CVPR 25] Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation
★ 282STream3R. Dynamic 3D Foundation Model using Causal Transformer. [ICLR 2026]
★ 393Splat-MOVER. Splat-MOVER: Multi-Stage, Open-Vocabulary Robotic Manipulation via Editable Gaussian Splatting
★ 48Affordance-R1. code for affordance-r1
★ 73Cross-View-AG. Official PyTorch Implementation of Learning Affordance Grounding from Exocentric Images, CVPR 2022
★ 75vrb. Python
★ 1373DAffordSplat. 3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians (ACM MM 25)
★ 79LS-Imagine. [ICLR 2025 Oral] PyTorch code for the paper "Open-World Reinforcement Learning over Long Short-Term Imagination"
★ 233dinov3. Reference PyTorch implementation and models for DINOv3
★ 11kawesome-representation-for-robotics.
★ 462TesserAct. ICCV 2025 | TesserAct: Learning 4D Embodied World Models
★ 403VR-Robo. [RA-L 2025] VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion
★ 188ReferSplat. [ICML2025 Oral] ReferSplat: Referring Segmentation in 3D Gaussian Splatting
★ 146infinigen. Infinite Photorealistic Worlds using Procedural Generation
★ 7.1kvipe. ViPE: Video Pose Engine for Geometric 3D Perception
★ 2.1kWorldGen. 🌍 WorldGen - Generate Any 3D Scene in Seconds
★ 2kCausVid. (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
★ 1.4kDifix3D. [CVPR 2025 Oral & Best Paper Finalist] Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
★ 1.2kFlexWorld. Official PyTorch implementation for "FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis".
★ 134FastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kPrior-Depth-Anything. Python
★ 520Gen3DSR. Gen3DSR: Generalizable 3D Scene Reconstruction via Divide and Conquer from a Single View, 3DV2025
★ 200Align3R. [CVPR 2025 Highlight] Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
★ 461VideoPainter. [SIGGRAPH2025] Official repo for paper "Any-length Video Inpainting and Editing with Plug-and-Play Context Control"
★ 626WorldScore. Official implementation for WorldScore: A Unified Evaluation Benchmark for World Generation
★ 302MoVieS. [CVPR 2026] Official implementation of "MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second".
★ 462SpaTrackerV2. [ICCV 2025] SpatialTrackerV2: 3D Point Tracking Made Easy
★ 985mega-sam. Code for the project "MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos"
★ 1.3kZPressor. [NeurIPS 2025] ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
★ 180DISCOVERSE-Real2Sim. Python
★ 62DISCOVERSE. Python
★ 489ScenePainter. [ICCV'25] ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
★ 37HoloMotion. HoloMotion: A Foundation Model for Whole-Body Humanoid Control
★ 605SAPIEN. SAPIEN Embodied AI Platform
★ 800robogs. Python
★ 554Self-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kdino_wm. Python
★ 536MetaScenes. Python
★ 68Real-Play. Code implementation for: From Virtual Games to Real-World Play
★ 48LangScene-X. [ICCV 2025] LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
★ 302VistaDream. [ICCV 2025] VistaDream: Sampling multiview consistent images for single-view scene reconstruction
★ 538coze-studio. An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before. Coze your way to AI Agent creation.
★ 21kFlash-Sculptor. Flash Sculptor: Modular 3D Worlds from Objects
★ 33APS-NeuS. APS-NeuS: Adaptive Planar and Skip Sampling for 3D Surface Reconstruction in High-Specular Scenes
★ 2Wan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kGaVS. [SIGGRAPH 2025] Official Repo for Paper - GaVS: 3D-Grounded Video Stabilization via Temporally-Consistent Local Reconstruction and Rendering
★ 293HunyuanWorld-1.0. Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
★ 2.9kNutWorld. Seeing World Dynamics in a Nutshell
★ 114omniphysgs. [ICLR 2025] OmniPhysGS: 3D Constitutive Gaussians for General Physics-based Dynamics Generation
★ 124Vid2Sim. [CVPR 2025] Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation
★ 43FreeGave. 🔥FreeGave in PyTorch (CVPR 2025)
★ 38Free4D. [ICCV 2025] Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
★ 2443DObjectReconstruction. 3D Object Reconstruction project is a workflow that takes a set of stereo images and camera info and outputs a textured mesh (i.e., .OBJ file). The purpose is to translate physical items into the digital world in a photorealistic way
★ 2264d-gaussian-splatting. [ICLR 2024] Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting
★ 1kMoGe. [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
★ 2.7kAmodal3R. Python
★ 135StableGS. Cuda
★ 9Pi3. [ICLR 2026] π^3: Permutation-Equivariant Visual Geometry Learning
★ 2.1kGLS. Geometry-aware 3D Language Gaussian Splatting (arXiv 2024)
★ 17VGGT-Long. Official implement of VGGT-Long
★ 885StreamVGGT. [ICLR 2026] Streaming 4D Visual Geometry Transformer
★ 945MaterialMVP. MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion
★ 206AnimateAnyMesh. [ICCV 2025] Official code for AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
★ 316articulate-anything. [ICLR 2025] Official implementation of Articulate-Anything
★ 205force-prompting. Official implementation of "Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals" (NeurIPS 2025)
★ 161VideoScene. [CVPR 2025 Highlight] VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
★ 353InstructScene. [ICLR 2024 spotlight] Official implementation of "InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior".
★ 152VoxelSplat. CVPR 2025: VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
★ 89vggt-blender. Blender addon for vggt 3D reconstruction
★ 78AnySplat. [SIGGRAPH Asia 2025 (ACM TOG)] AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
★ 899jepa. PyTorch code and models for V-JEPA self-supervised learning from video.
★ 4.1kDreamCube. [ICCV 2025] Official implementation of the paper "DreamCube: 3D Panorama Generation via Multi-plane Synchronization".
★ 181SD-T2I-360PanoImage. repository for 360 panorama image generation based on Stable Diffusion
★ 324Pano2Room. Pytorch implementation of Pano2Room (SIGGRAPH Asia 2024)
★ 152RAPO. [CVPR 2025] The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
★ 106on-the-fly-nvs. Official implementation of On-the-fly Reconstruction for Large-Scale Novel View Synthesis from Unposed Images. A. Meuleman, I. Shah, A. Lanvin, B. Kerbl, G. Drettakis, ACM TOG (proc. SIGGRAPH) 2025
★ 572umeyama-python. Python Umeyama's Algorithm - Transformation between two sets of points
★ 36VisAnything. Python
★ 30Paper2Poster. [NeurIPS 2025] Open-source Multi-agent Poster Generation from Papers
★ 3.9kzerohsi. JavaScript
★ 8HowToCook. Programmer's guide about how to cook at home.
★ 101kglomap. [DEPRECATED] GLOMAP - Global Structured-from-Motion Revisited
★ 2.4kGenFusion. [CVPR 2025] GenFusion: Closing the Loop between Reconstruction and Generation via Videos
★ 176pointrix. A differentiable point-based rendering framework.
★ 218FDS. [ICLR 2025] Flow Distillation Sampling: Regularizing 3D Gaussians with Pre-trained Matching Priors
★ 65DiffDIS. High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity (ICLR2025)
★ 59