This is your work, valued
HNR-VLN. Official implementation of Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation (CVPR'24 Highlight).
★ 109GridMM. Official implementation of GridMM: Grid Memory Map for Vision-and-Language Navigation (ICCV'23).
★ 106Dynam3D. Official implementation of "Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation" (NeurIPS'25 Oral)
★ 89Sim2Real-VLN-3DFF. Official implementation of Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation (CoRL'24).
★ 80NavRAG. Official implementation of "NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM" (ACL'25 Findings)
★ 60g3D-LF. Official implementation of "g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks" (CVPR'25).
★ 57Super-resolution-SR-CUDA. High performance SRMD implementation using CUDA.
★ 29D3D-VLP.
★ 17sage. Official Code Release of SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
★ 375AnyRecon. [SIGGRAPH Asia 2026] AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
★ 389GeoWorld. [ECCV 2026] GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation
★ 14worldmesh. [ECCV 2026] WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
★ 145Spatia. [CVPR2026] Long-horizon, spatially consistent video generation enabled by persistent 3D scene point clouds and dynamic-static disentanglement.
★ 223PixWorld. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
★ 235FlexWorld. Official PyTorch implementation for "FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis".
★ 134LatentSpatialMemory. Latent Spatial Memory for Video World Models
★ 287GaussianGPT. [ECCV'26] Our method creates 3D Gaussian scenes completely autoregressively, allowing for generation, completion, and outpainting with the same model.
★ 301G4Splat. Official implementation of ICLR26 paper "G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior"
★ 146robot-PointAct. Official implementation of PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction (RSS'26).
★ 14Moebius. [ECCV 2026] Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
★ 508Prompt-Inpaint. An end-to-end pipeline for text-prompt detection, SAM3 segmentation, and background inpainting. 文本提示检测 + SAM3 分割 + 背景修复(inpaint)的端到端管线
★ 2drifting-model. Personal PyTorch implementation of "Generative Modeling via Drifting" with Claude
★ 241DriftingModel. PyTorch implementation of Drifting Models by Kaiming He et al.
★ 20MV-SAM3D. SAM 3D Objects with Multi-view Images
★ 512motrixsim-docs. A high-performance physics simulation engine designed for multibody dynamics and robotics simulation.
★ 147gs_playground. Jupyter Notebook
★ 453lingbot-vision. Self-supervised learning for spatial perception
★ 887Argus. [ECCV 2026] Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
★ 65ArtiFixer. Python
★ 583robocasa. RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
★ 1.6kgenie_sim. Simulation Platform from AgiBot
★ 1.3krecgen. Jupyter Notebook
★ 203stretch_mujoco. This library provides a simulation stack for Stretch, built on MuJoCo.
★ 63HomeWorld.
★ 77GN0. The official Implementation of GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation
★ 41image-blaster. An image-to-world skillset for Claude.
★ 4.8klingbot-map. A feed-forward 3D foundation model for reconstructing scenes from streaming data
★ 16kmvdust3r. Open source impl of **MV-DUSt3R+ Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds** from Meta Reality Labs. Project page https://mv-dust3rp.github.io/
★ 601AnySplat. [SIGGRAPH Asia 2025 (ACM TOG)] AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
★ 901splatter360. [CVPR 2025] Splatter-360: Generalizable 360 Gaussian Splatting for Wide-baseline Panoramic Images
★ 171PanSplat. 🍳 [CVPR'25] PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting
★ 232molmospaces. An end-to-end open ecosystem for robot learning
★ 423AgentVLN.
★ 180PointWorld. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
★ 469pMF. Official Implementation of pMF https://arxiv.org/abs/2601.22158
★ 270lingbot-vla. A Pragmatic VLA Foundation Model
★ 1.7kimeanflow. Official Implementation of iMF https://arxiv.org/abs/2512.02012
★ 335meanflow. JAX implementation of MeanFlow
★ 602linkerhand-python-sdk. Linkerhand Python SDK
★ 44diff-gaussian-feature-rasterization. Differentiable Gaussians' Features Rasterization Submodule
★ 1vdpm. Official implementation of Video-DPM
★ 242DexTeleGym. After modifications on OpenTelevision, Quest 3 can be used for teleoperation of the Franka Panda robotic arm and Inspire Hand in Isaac Gym.
★ 63inspire_hands. working with the inspire robot hands (RH56 series)
★ 63inspire_demos. Python
★ 8Action100M. A Large-scale Video Action Dataset
★ 483FeatUp. Official code for "FeatUp: A Model-Agnostic Frameworkfor Features at Any Resolution" ICLR 2024
★ 1.7klightning-grasp. The Lightning Grasp System
★ 204AstraNav-World. Official implementation of [AstraNav-World: World Model for Foresight Control and Consistency]
★ 99OmniNav. 【ICLR 2026】 Official implementation of [OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation]
★ 202docker-selkies-egl-desktop. KDE Plasma Desktop container designed for Kubernetes, supporting OpenGL EGL and GLX, Vulkan, and Wine/Proton for NVIDIA GPUs through WebRTC and HTML5, providing an open-source remote cloud/HPC graphics or game streaming platform.
★ 344IAmGoodNavigator. Python
★ 91RealSee3D. RealSee3D: A multi-view RGB-D dataset combining real-world captures and procedurally generated scenes, with extensible annotations for diverse 3D vision research.
★ 282SID-VLN. [ECCV 2026] Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale
★ 14Depth-Anything-3. Depth Anything 3
★ 6kDiT360. [CVPR 2026] Official implementation of "DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training".
★ 279JanusVLN. [ICLR2026] Official implementation for "JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation"
★ 575se3ds. This repository hosts the code for our paper, "Simple and Effective Synthesis of Indoor 3D Scenes".
★ 42world-in-world. "World Models in a Closed-Loop World" (ICLR'26 Oral)
★ 181Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kgrounded-rl. Python
★ 133TTT3R. [ICLR 2026] A simple state update rule to enhance length generalization for CUT3R
★ 714MTU3D. Python
★ 266Nav-R1. Nav-R1: Reasoning and Navigation in Embodied Scenes
★ 1303DLLM-Mem.
★ 27GraspMolmo. Code and website for "GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation"
★ 46Holodeck. CVPR 2024: Language Guided Generation of 3D Embodied AI Environments.
★ 561Vid2Sim. [CVPR 25] Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation
★ 283SceneCompleter. SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
★ 36thortils. Utility functions when working with Ai2-THOR. Try to do one thing once.
★ 57A0. Python
★ 76RynnEC. RynnEC: Bringing MLLMs into Embodied World
★ 391scaling_on_scales. When do we not need larger vision models?
★ 419dinov3. Reference PyTorch implementation and models for DINOv3
★ 11kVLN_CLASH. This is the official repository for VLN-CLASH.
★ 27DifNav. This is the source code to paper “DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation”.
★ 34Being-H. Being-H is BeingBeyond's family of human-centric embodied foundation models.
★ 1.1kVPN. Official implementation of "VPN: Visual Prompt Navigation"
★ 8SmartCLIP. SmartCLIP: A training method to improve CLIP with both short and long texts
★ 43habitat-sim. A flexible, high-performance 3D simulator for Embodied AI research.
★ 3.8kInternNav. InternRobotics' open platform for building generalized navigation foundation models.
★ 1kstretch_isaacsim.
★ 7LOVMM. LOVMM
★ 12Pi3. [ICLR 2026] π^3: Permutation-Equivariant Visual Geometry Learning
★ 2.1kobjaverse-xl. 🪐 Objaverse-XL is a Universe of 10M+ 3D Objects. Contains API Scripts for Downloading and Processing!
★ 1.3kSpaTrackerV2. [ICCV 2025] SpatialTrackerV2: 3D Point Tracking Made Easy
★ 986MoMaKitchen. [ICCV 2025] MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
★ 55StreamVLN. [ICRA 2026] Official implementation of the paper: "StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling"
★ 565NaVILA. [RSS'25] This repository is the implementation of "NaVILA: Legged Robot Vision-Language-Action Model for Navigation"
★ 677YOTO. [TPAMI2026 / RSS2025] Code for my paper "You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations"
★ 147LAPA. [ICLR 2025] LAPA: Latent Action Pretraining from Videos
★ 562mobipi. Official implementation for Mobi-π.
★ 128NavMorph. Official implementation of NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments (ICCV'25).
★ 87TrackVLA. [CoRL 2025] Repository relating to "TrackVLA: Embodied Visual Tracking in the Wild"
★ 420NavDP. [ICRA 2026] NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
★ 728Object2HabitatMap. Awesome habitat top down map work 🤩
★ 35Uni-NaVid. [RSS 2025] Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.
★ 333instant_policy.
★ 3brs-algo. Official Algorithm Codebase for the Paper "BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities"
★ 171transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kIsaac-GR00T. NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7.7kHA-VLN. Official implementation for "HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions".
★ 396stretch_community. Portal for resources for the Stretch community
★ 13TinySAM. [AAAI 2025] Official PyTorch implementation of "TinySAM: Pushing the Envelope for Efficient Segment Anything Model"
★ 555EMMOE. EMMOE: A Comprehensive Benchmark for Embodied Mobile Manipulation in Open Environments
★ 28MSR3D. [NeurIPS 2024] MSR3D: Multimodal Situated Reasoning in 3D Scenes
★ 77procthor. 🏘️ Scaling Embodied AI by Procedurally Generating Interactive 3D Houses
★ 461multiscan. [NeurIPS 2022] MultiScan: Scalable RGBD scanning for 3D environments with articulated objects
★ 151SoFar. [NeurIPS 2025 Spotlight] SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
★ 246OpenFly-Platform. C++
★ 350academicpages.github.io. Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.
★ 17kUni-NaVid. JavaScript
★ 13FLaRe. [ICRA 25] FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
★ 49mpao. Python
★ 7stretch-open. Python
★ 21SegmentAnything3D. [ICCV'23 Workshop] SAM3D: Segment Anything in 3D Scenes
★ 1.4kFastSAM. Fast Segment Anything
★ 8.4kanygrasp_sdk. Python
★ 953ok-robot. An open, modular framework for zero-shot, language conditioned pick-and-drop tasks in arbitrary homes.
★ 609spoc-robot-training. SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
★ 156RVT. Official Code for RVT-2 and RVT
★ 408ManiSkill. Manipulation Skill Framework, an open source GPU parallelized robotics simulator and benchmark
★ 3.2kstretch_urdf. URDFs for the Stretch mobile manipulators from Hello Robot Inc.
★ 15EmbodiedScan. [CVPR 2024 & NeurIPS 2024] EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
★ 672openpi. Python
★ 13kVideo-3D-LLM. [CVPR 2025] The code for paper ''Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding''.
★ 220VLN-CE. Vision-and-Language Navigation in Continuous Environments using Habitat
★ 844rectified-flow. 从零手搓Flow Matching(Rectified Flow)
★ 637ESAM. [ICLR 2025, Oral] EmbodiedSAM: Online Segment Any 3D Thing in Real Time
★ 634stretch_visual_servoing. Example code for visual servoing using Stretch 3's gripper camera
★ 21dobb-e. Dobb·E: An open-source, general framework for learning household robotic manipulation
★ 621harmonic_mobile_manipulation. Python
★ 17legged-loco. Low-level locomotion policy training in Isaac Lab
★ 447NaVILA-Bench. Vision-Language Navigation Benchmark in Isaac Lab
★ 331NaVid-VLN-CE. [RSS 2024 & RSS 2025] VLN-CE evaluation code of NaVid and Uni-NaVid
★ 438mshab. A Benchmark for Low-Level Manipulation in Home Rearrangement Tasks
★ 196UnSAM. [NeurIPS 2024] Code release for "Segment Anything without Supervision"
★ 503LightRAG. [EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation"
★ 38kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kforcesight. Given an RGBD image and a text prompt, ForceSight produces visual-force goals for a robot, enabling mobile manipulation in unseen environments with unseen object instances.
★ 25open-eqa. OpenEQA Embodied Question Answering in the Era of Foundation Models
★ 365Embodied_RAG. Python
★ 55CLIP-SAM. Experiment on combining CLIP with SAM to do open-vocabulary image segmentation.
★ 389stretch_ai. Python
★ 230InfiniteWorld. Python
★ 86VLN-SRDF. Official implementation of: Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
★ 35SAME. [ICCV 2025] Official implementation of SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
★ 40point_lio_unilidar. Point-LIO algorithm for Unitree LiDAR products.
★ 505vrb. Python
★ 137MapGPT. [ACL 24] The official implementation of MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation.
★ 134open_x_embodiment. Jupyter Notebook
★ 2kLLaVA-3D. [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D World
★ 388OpenING. Official Implementation of OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
★ 252AdaVLN. IsaacSim Extension for Dynamic Objects in Matterport3D Environments for AdaVLN research
★ 68feature-splatting-inria. Feature splatting based on INRIA GS rasterizer
★ 106concept-graphs. Official code release for ConceptGraphs
★ 915VLN-Game. A new zero-shot framework to explore and search for the language descriptive targets in unknown environment based on Large Vision Language Model.
★ 78PoliFormer. PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
★ 107partnr-planner. A repository accompanying the PARTNR benchmark for using Large Planning Models (LPMs) to solve Human-Robot Collaboration or Robot Instruction Following tasks in the Habitat simulator.
★ 379SummaryOfLoanSuspension. 全国各省市停贷通知汇总
★ 20kPIVOT-R. [NeurIPS 2024] PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
★ 49graphrag. A modular graph-based Retrieval-Augmented Generation (RAG) system
★ 35krsrd. Code for "Robot See Robot Do" presented at CoRL 2024!
★ 163learning-language-navigation. Official repository for LeLaN training and inference code
★ 142Chat-Scene. [NeurIPS 2024 & TPAMI 2026] Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
★ 216RoboticsDiffusionTransformer. RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
★ 1.8kDigitalTwinArt. Python
★ 71ovon. Open Vocabulary Object Navigation
★ 140carbonyl. Chromium running inside your terminal
★ 19kbypy. Python client for Baidu Yun (Personal Cloud Storage) 百度云/百度网盘Python客户端
★ 8.6kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kopenvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7ke2e-NeRF-nav.
★ 1InstructNav. Python
★ 210open_clip. An open source implementation of CLIP.
★ 14khabitat_tools. python tools to work with habitat-sim environment.
★ 43ScanReason. [ECCV 2024] Empowering 3D Visual Grounding with Reasoning Capabilities
★ 85VLA-3D. Python
★ 113vil3dref. Official implementation of Language Conditioned Spatial Relation Reasoning for 3D Object Grounding (NeurIPS'22).
★ 67SQA3D. [ICLR 2023] SQA3D for embodied scene understanding and reasoning
★ 170feature-splatting. Python
★ 160Long-CLIP. [ECCV 2024] official code for "Long-CLIP: Unlocking the Long-Text Capability of CLIP"
★ 901PQ3D. Official implementation of the paper "Unifying 3D Vision-Language Understanding via Promptable Queries"
★ 86Director3D. Code for "Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text" (NeurIPS 2024).
★ 381BEVInstructor. [ECCV24] Navigation Instruction Generation with BEV Perception and Large Language Models
★ 31GaussNav. PyTorch implementation of paper: GaussNav: Gaussian Splatting for Visual Navigation
★ 218robot_sugar. Official implementation of "SUGAR: Pre-training 3D Visual Representations for Robotics" (CVPR'24).
★ 46BunnyVisionPro. Bimanual Dexterous Teleoperation with Real-Time Retargeting using VisionPro
★ 355TeleVision. [CoRL 2024] Open-TeleVision: Teleoperation with Immersive Active Visual Feedback
★ 1.3kVLN-PRET. Jupyter Notebook
★ 23InternUtopia. A simulation platform for versatile Embodied AI research and developments.
★ 1.3kNavGPT-2. [ECCV 2024] Official implementation of NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
★ 245VLN-MAGIC. This is the official repository for MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation Learning towards Efficient Vision-and-Language Navigation
★ 17HA3D_simulator. Official implementation of Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions (NeurIPS DB Track'24 Spotlight).
★ 58LangSplat. Official implementation of the paper "LangSplat: 3D Language Gaussian Splatting" [CVPR2024 Highlight]
★ 1.1kDiaLoc. Official implementation of DiaLoc: An Iterative Approach to Embodied Dialog Localization (CVPR 2024)
★ 7SceneVerse. Official implementation of ECCV24 paper "SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding"
★ 289vlfm. The repository provides code associated with the paper VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation (ICRA 2024)
★ 784FreeSplat. Official implementation of NeurIPS 2024 paper: "FreeSplat: Generalizable 3D Gaussian Splatting Towards Free-View Synthesis of Indoor Scenes"
★ 167