This is your work, valued
Embodied AI Researcher @OpenRobotLab, Shanghai AI Lab
SceneMover. Project of Siggraph Asia 2020 paper: Scene Mover: Automatic Move Planning for Scene Arrangement by Deep Reinforcement Learning
★ 98PSVH-3d-reconstruction. Deep Single-View 3D Object Reconstruction with Visual Hull Embedding
★ 48Active_VLN. The repository of ECCV 2020 paper `Active Visual Information Gathering for Vision-Language Navigation`
★ 44SSM-VLN. Code and Data for our CVPR 2021 paper "Structured Scene Memory for Vision-Language Navigation"
★ 43CCC-VLN. Repository of our accepted CVPR2022 paper "Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation"
★ 28VXN. Repository of our accepted NeurIPS-2022 paper "Towards Versatile Embodied Navigation"
★ 22Dreamwalker. Implementation of our ICCV 2023 paper DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation
★ 20MelodyMaster. Microsoft Hackthon 2017 Award of Most Creative
★ 5Polygon_Intersection. Algorithm to calculate the intersected area of polygons
★ 2Literature-Search. High level understanding
★ 1MICE. MICE Imputation implementation using scikit learn.
★ 1AlphaPose. Real-Time and Accurate Multi-Person Pose Estimation&Tracking System
★ 1ABot-Navigation. Python
★ 205lingbot-world-v2. Infinite Worlds with Versatile Interactions
★ 1.4klingbot-depth. Masked Depth Modeling for Spatial Perception
★ 1.5kMotus. Official code of Motus: A Unified Latent Action World Model
★ 1.2kEBench. Elemental Diagnosis of Generalist Mobile Manipulation Policies
★ 122hermes-agent. The agent that grows with you
★ 223kSIM1. Official implementation of "SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds "
★ 165FastWAM. Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
★ 1.2kInternDataEngine. InternDataEngine: Pioneering High-Fidelity Synthetic Data Generator for Robotic Manipulation
★ 120DrClaw. Python
★ 170openclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kUltraDexGrasp. [ICRA 2026] UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
★ 84Xiaomi-Robotics-0. Python
★ 627DreamDojo. Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
★ 1kawesome-claude-skills. A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows
★ 71klehome-challenge. Python
★ 146giga-brain-0. GigaBrain-0: A World Model-Powered Vision-Language-Action Model
★ 2.6kcosmos-policy. Cosmos Policy
★ 842dreamzero. Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
★ 2.5kRDT2. Official code of RDT 2
★ 797lingbot-vla. A Pragmatic VLA Foundation Model
★ 1.7klingbot-va. [RSS 2026] Causal video-action world model for generalist robot control
★ 1.7kurdf-loaders. URDF Loaders for Unity and THREE.js with example ATHLETE URDF Files open sourced from NASA JPL
★ 807dvc. 🦉 Data Versioning and ML Experiments
★ 16kVLAC. VLAC: A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
★ 321large-video-planner. Python
★ 256VL-LN. VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs
★ 58Action100M. A Large-scale Video Action Dataset
★ 483RLinf. RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
★ 4.3kspirit-v1.5. Spirit-v1.5: A Robotic Foundation Model by Spirit AI
★ 631Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kQwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kInternVLA-A-series. InternVLA-A1: Unifying Understanding, Generation, and Action for Robotic Manipulation
★ 518behavior-1k-solution. 1st place solution of 2025 BEHAVIOR Challenge
★ 317openpi-comet. Team Comet's 2025 BEHAVIOR Challenge Codebase
★ 262IsaacLab-Arena. Isaac Lab - Arena is a robotics simulation framework that enhances NVIDIA Isaac Lab by providing a composable, scalable system for creating diverse simulation environments and evaluating robot learning policies. The framework enables developers to rapidly prototype and test robotic tasks with various robot embodiments, objects, and environments.
★ 505GraspGen. Official repo for GraspGen: A Diffusion-based Framework for 6-DOF Grasping
★ 532MeshCoder. Jupyter Notebook
★ 500PhysX-Anything. PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image (CVPR 2026)
★ 911X-VLA. [ICLR 2026] The offical Implementation of "Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model"
★ 697starVLA. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
★ 3.3kmjlab. Isaac Lab API, powered by MuJoCo-Warp, for RL and robotics research
★ 2.8kdexbotic. Dexbotic: Open-Source Vision-Language-Action Toolbox
★ 1.3kF1-VLA. F1: A Vision Language Action Model Bridging Understanding and Generation to Actions
★ 200InternVLA-M1. InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
★ 418libuipc. A Modern C++20 Library of Unified Incremental Potential Contact.
★ 307Grounded-SAM-2. Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2
★ 3.7kGrounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kInternNav. InternRobotics' open platform for building generalized navigation foundation models.
★ 1kInternSR. InternRobotics' open-source toolbox for vision-based embodied spatial intelligence.
★ 49InternHumanoid. A versatile, all-in-one toolbox for whole-body humanoid robot control.
★ 185InternManip. An All-in-one robot manipulation learning suite for policy models training and evaluation on various datasets and benchmarks.
★ 175lerobot-sim2real. LeRobot sim2real code. Train in fast simulation and deploy visual policies zero shot to the real world
★ 389OST-Bench. [NeurIPS 2025] OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
★ 80StreamVLN. [ICRA 2026] Official implementation of the paper: "StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling"
★ 565Lightwheel-simready-asset. Open-source 3D digital assets for simulation and robot training. Free for non-commercial use.
★ 171motphys-rigidbody-unity-sdk. Motphys Rigidbody Physics SDK for Unity
★ 29newton. An open-source, GPU-accelerated physics simulation engine built upon NVIDIA Warp, specifically targeting roboticists and simulation researchers.
★ 5.3kgemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kpgnd. Code implementation for RSS 2025 paper Particle-Grid Neural Dynamics for Learning Deformable Object Models from RGB-D Videos
★ 87code-server. VS Code in the browser
★ 79kCronusVLA. [AAAI26 oral] CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
★ 1123dgrut. Ray tracing and hybrid rasterization of Gaussian particles
★ 2.3kAwesome-World-Models. A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related websites.
★ 1.9kGR00T-Dreams. DreamGen: Nvidia GEAR Lab's initiative to solve the robotics data problem using world models
★ 592Tencent-XR-3DGen. Python
★ 356vggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kIsaacSim. NVIDIA Isaac Sim™ is an open-source application on NVIDIA Omniverse for developing, simulating, and testing AI-driven robots in realistic virtual environments.
★ 3.8kGenManip. [CVPR 2025] Official implementation of "GenManip: LLM-driven Simulation for Generalizable Instruction-Following Manipulation"
★ 170real-time-chunking-kinetix. Simulated experiments for "Real-Time Execution of Action Chunking Flow Policies".
★ 553Direct3D-S2. [NeurIPS 2025] Direct3D‑S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
★ 1.3kHSMR. [CVPR25 Oral (Top 3.3%)] Official code for paper "Reconstructing Humans with a Biomechanically Accurate Skeleton".
★ 632MMaDA. MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
★ 1.7kIsaac-GR00T. NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7.7kopenvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7kNavDP. [ICRA 2026] NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
★ 725RoboVLMs. Python
★ 474TAPIP3D. TAPIP3D: Tracking Any Point in Persistent 3D Geometry
★ 412HoST. [RSS 2025 Best Systems Paper Finalist] 💐Official implementation of "Learning Humanoid Standing-up Control across Diverse Postures"
★ 602Infinite-Mobility. C++
★ 195genie_sim. Simulation Platform from AgiBot
★ 1.3kDiffTactile. [ICLR 2024] DiffTactile: A Physics-based Differentiable Tactile Simulator for Contact-rich Robotic Manipulation
★ 315owl. 🦉 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
★ 20kRoboVerse. RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
★ 1.8kTokenHSI. [CVPR 2025 Oral] TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
★ 558Aether. [ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling
★ 604mmc4. MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.
★ 954warp. A Python framework for GPU-accelerated simulation, robotics, and machine learning.
★ 6.9kopenvla-oft. Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
★ 1.3kstreamlit. Streamlit — A faster way to build and share data apps.
★ 45kopenpi. Python
★ 13kSpatialVLA. 🔥 SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.
★ 711any4lerobot. 🎁 A collection of utilities for LeRobot.
★ 1.1kcamel. 🐫 CAMEL: The first and the best multi-agent framework. Finding the Scaling Law of Agents. https://www.camel-ai.org
★ 18kInternUtopia. A simulation platform for versatile Embodied AI research and developments.
★ 1.3kSoFar. [NeurIPS 2025 Spotlight] SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
★ 246RoboTwin. [ICML 2026] RoboTwin 2.0 Offical Repo
★ 2.7kfractalgen. PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
★ 1.2ksimple_GRPO. A very simple GRPO implement for reproducing r1-like LLM thinking.
★ 1.7kOpenHomie. Open-sourced code for "HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit".
★ 596MoGe. [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
★ 2.7kHumanoidVerse. Python
★ 462open-r1. Fully open reproduction of DeepSeek-R1
★ 26kDeepSeek-R1.
★ 92knerfies.github.io. JavaScript
★ 4.3kDISCOVERSE. Python
★ 490TRELLIS. Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
★ 13kPoliFormer. PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
★ 107ProtoMotions. ProtoMotions is a GPU-accelerated simulation and learning framework for training physically simulated digital humans and humanoid robots.
★ 2.2kRoboticsDiffusionTransformer. RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
★ 1.8kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kprocthor. 🏘️ Scaling Embodied AI by Procedurally Generating Interactive 3D Houses
★ 461Seer. [ICLR 2025 Oral] Seer: Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
★ 310genesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kDIDEX. Python
★ 29act-plus-plus. Imitation learning algorithms with Co-training for Mobile ALOHA: ACT, Diffusion Policy, VINN
★ 3.6kSimGen. Simulator-conditioned Driving Scene Generation
★ 139depthcues. [CVPR 2025] "DepthCues: Evaluating Monocular Depth Perception in Large Vision Models", Duolikun Danier, Mehmet Aygün, Changjian Li, Hakan Bilen, Oisin Mac Aodha
★ 23VAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kSimplerEnv. Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Google Robot, WidowX+Bridge) (CoRL 2024)
★ 1.1khamer. HaMeR: Reconstructing Hands in 3D with Transformers
★ 1.1kDeepRL. Deep Reinforcement Learning Lab, a platform designed to make DRL technology and fun for everyone
★ 2.6kawesome-3D-gaussian-splatting. Curated list of papers and resources focused on 3D Gaussian Splatting, intended to keep pace with the anticipated surge of research in the coming months.
★ 8.8kOminiControl. [ICCV 2025 Highlight] OminiControl: Minimal and Universal Control for Diffusion Transformer
★ 1.9ksamurai. Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
★ 7.1kWHAM. Python
★ 1.1kWiLoR. WiLoR: End-to-end 3D hand localization and reconstruction in-the-wild
★ 6133D-Diffusion-Policy. [RSS 2024] 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
★ 1.4kPhysX. NVIDIA PhysX SDK
★ 4.7kawesome-robotics-libraries. :sunglasses: A curated list of robotics libraries and software
★ 3kAwesome-Robotics-Manipulation. A comprehensive list of papers about Robot Manipulation, including papers, codes, and related websites.
★ 1.1kPLATO. Python
★ 3point2mesh. Reconstruct Watertight Meshes from Point Clouds [SIGGRAPH 2020]
★ 1.2kDepth-Anything. [CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation
★ 8.2kawesome-simulation. Resources for Physics based simulation in Computer Graphics 图形学中物理模拟的资源整理
★ 393Awesome-LLM-Strawberry. A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.
★ 6.9kSAM2Point. The Most Faithful Implementation of Segment Anything (SAM) in 3D
★ 359co-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5ktapnet. Tracking Any Point (TAP)
★ 2kLightZero. [NeurIPS 2023 Spotlight] LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios (awesome MCTS)
★ 1.6ksapiens. High-resolution models for human tasks.
★ 5.4kLLM-Agent-Paper-List. The paper list of the 86-page SCIS cover paper "The Rise and Potential of Large Language Model Based Agents: A Survey" by Zhiheng Xi et al.
★ 8.2ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kmink. Python inverse kinematics based on MuJoCo.
★ 1.5kHamba. [NeurIPS 2024] Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mamba
★ 186LEGENT. Open Platform for Embodied Agents
★ 342awesome-diffusion-model-in-rl. A curated list of Diffusion Model in RL resources (continually updated)
★ 1.6kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kdex-urdf. Python
★ 364mujoco_menagerie. A collection of high-quality models for the MuJoCo physics engine, curated by Google DeepMind.
★ 3.8krobot_descriptions.py. Access 190+ robot descriptions from the main Python robotics frameworks
★ 799ORB_SLAM2. Real-Time SLAM for Monocular, Stereo and RGB-D Cameras, with Loop Detection and Relocalization Capabilities
★ 10kPRS-Trial-Version. Trial version for prs platform (python project). Please note that the complete experience requires downloading the Unity resource.
★ 10awesome-llm-powered-agent. Awesome things about LLM-powered agents. Papers / Repos / Blogs / ...
★ 2.3kIsaacLab. Unified framework for robot learning built on NVIDIA Isaac Sim
★ 7.8kredoc. 📘 OpenAPI/Swagger-generated API Reference Documentation
★ 26kviser. Web-based 3D visualization in Python
★ 2.7kEmbodiedScan. [CVPR 2024 & NeurIPS 2024] EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
★ 672kaolin. A PyTorch Library for Accelerating 3D Deep Learning Research
★ 5.1ktripy. Simple polygon triangulation algorithms in pure python
★ 65Awesome-Embodied-AI. A curated list of awesome papers on Embodied AI and related research/industry-driven resources.
★ 527langsuite. Official Repo of LangSuitE
★ 85autogen. A programming framework for agentic AI
★ 60kHierarchical-Localization. Visual localization made easy with hloc
★ 4.2kzero123plus. Code repository for Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model.
★ 2.1kEureka. Official Repository for "Eureka: Human-Level Reward Design via Coding Large Language Models" (ICLR 2024)
★ 3.2kMetaCLIP. NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024
★ 1.9kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kAwesome-LLM-Robotics. A comprehensive list of papers using large language/multi-modal models for Robotics/RL, including papers, codes, and related websites
★ 4.4kpartnet_dataset. PartNet Dataset Official Release Repo
★ 422Awesome-Open-Vocabulary-Semantic-Segmentation. A curated publication list on open vocabulary semantic segmentation and related area (e.g. zero-shot semantic segmentation) resources..
★ 893Octopus. [ECCV2024] 🐙Octopus, an embodied vision-language model trained with RLEF, emerging superior in embodied visual planning and programming.
★ 301CogVLM. a state-of-the-art-level open visual language model | 多模态预训练模型
★ 6.7kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kfiftyone. Refine high-quality datasets and visual AI models
★ 11klanguage-refer. Python
★ 27LongLoRA. Code and documents of LongLoRA and LongAlpaca (ICLR 2024 Oral)
★ 2.7kreferit3d. Code accompanying our ECCV-2020 paper on 3D Neural Listeners.
★ 141Awesome-Agents-Research.
★ 121depRL. Repository for our ICLR 2023 paper: DEP-RL: Embodied Exploration for Reinforcement Learning in Overactuated and Musculoskeletal Systems
★ 312VLN-BEVBert. [ICCV 2023] Official repo of "BEVBert: Multimodal Map Pre-training for Language-guided Navigation"
★ 260openinterpreter. A coding agent for open models like Kimi K3
★ 67kPointLLM. [ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
★ 1kspear. SPEAR: A Simulator for Photorealistic Embodied AI Research
★ 543AgentVerse. 🤖 AgentVerse 🪐 is designed to facilitate the deployment of multiple LLM-based agents in various applications, which primarily provides two frameworks: task-solving and simulation
★ 5.1kbetterbib. :green_book: Command-line tools for bibliographies.
★ 839Dreamwalker. Implementation of our ICCV 2023 paper DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation
★ 20sim-web-visualizer. Web Based Visualizer for Simulation Environments
★ 419MetaGPT. 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
★ 70ksd-webui-controlnet. WebUI extension for ControlNet
★ 18kThe-Art-of-Linear-Algebra. Graphic notes on Gilbert Strang's "Linear Algebra for Everyone"
★ 22kLISA. Project Page for "LISA: Reasoning Segmentation via Large Language Model"
★ 2.7kgenerative_agents. Generative Agents: Interactive Simulacra of Human Behavior
★ 22kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kKandinsky-2. Kandinsky 2 — multilingual text2image latent diffusion model
★ 2.8kScaleVLN. [ICCV 2023 Oral]: Scaling Data Generation in Vision-and-Language Navigation
★ 225pddlstream. PDDLStream: Integrating Symbolic Planners and Blackbox Samplers
★ 4833D-LLM. Code for 3D-LLM: Injecting the 3D World into Large Language Models
★ 1.2kLLMAgentPapers. Must-read Papers on LLM Agents.
★ 3.1k