This is your work, valued
Master Student in Robotics at Carnegie Mellon University
Holistic-Robot-Pose-Estimation. [ECCV 2024] PyTorch implementation of "Real-time Holistic Robot Pose Estimation with Unknown States"
★ 40Holistic-Robot-Pose. HTML
★ 1Awesome-VLA-Post-Training. A collection of vision-language-action model post-training methods.
★ 233awesome-VLA-WAM. World Action Model, VLA failure detection/correction, Efficient VLA
★ 12rlpd. Python
★ 412HaWoR. HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
★ 310Humanoid-self-other-distincton-model.
★ 5HumanEgo. HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
★ 467GazeVLA. Python
★ 39cap-x. A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
★ 683simtoolreal. Official implementation of SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation
★ 336VGHuman. Official PyTorch implementation of "Visually-grounded Humanoid Agents"
★ 49Awesome-Touch. Tactile Sensing • Data Collection • IL/RL/VLA/WM • Manipulation • Simulation • Open Source
★ 751Awesome-RL-VLA. A Survey on Reinforcement Learning of Vision-Language-Action Models for Robotic Manipulation
★ 812openpi_pytorch. Pytorch PI-zero and PI-zero-fast. Adapted from LeRobot
★ 206pyroki. A Modular Toolkit for Robot Kinematic Optimization
★ 1.7kVisualMimic. [arXiv 2025] VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
★ 298holosoma. Python
★ 1.6kGMR. [ICRA 2026] GMR: General Motion Retargeting. Retarget human motions into diverse humanoid robots in real time on CPU. Retargeter for TWIST.
★ 2.5kembodied-ai-start. [PKU EPIC Lab] 面向小白的具身智能入门指南
★ 1.1kArticuBot. Official repository for RSS 25 paper: ArticuBot. Project: https://articubot.github.io/
★ 41phosphobot. Control AI robots. Community-driven UI middleware for controlling robots, recording datasets, training action models. Compatible with SO-100 and SO-101
★ 388Open-Teach. A Versatile Teleoperation framework for Robotic Manipulation using Meta Quest3
★ 417oculus_reader. C++
★ 166loco-mujoco. Imitation learning benchmark focusing on complex locomotion tasks using MuJoCo.
★ 1.4kHSMR. [CVPR25 Oral (Top 3.3%)] Official code for paper "Reconstructing Humans with a Biomechanically Accurate Skeleton".
★ 632Improved-3D-Diffusion-Policy. [IROS 2025] Generalizable Humanoid Manipulation with 3D Diffusion Policies. Part 1: Train & Deploy of iDP3
★ 552awesome-embodied-vla-va-vln. A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
★ 3.4kPHC. Official Implementation of the ICCV 2023 paper: Perpetual Humanoid Control for Real-time Simulated Avatars
★ 1.3kGenesisEnvs. Genesis Reinforcement Learning Environments
★ 137recognize-anything. Open-source and strong foundation image recognition models.
★ 3.7kRL4VLA. Python
★ 279vlarl. Single-file implementation to advance vision-language-action (VLA) models with reinforcement learning.
★ 447SimpleVLA-RL. [ICLR 2026] SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
★ 1.8kSelfSimRobot. Python
★ 28UP-VLA. Official PyTorch implementation for ICML 2025 paper: UP-VLA.
★ 61MathVista. MathVista: data, code, and evaluation for Mathematical Reasoning in Visual Contexts
★ 367EmbodiedBench. [ICML 2025 Oral] Official repo of EmbodiedBench, a comprehensive benchmark designed to evaluate MLLMs as embodied agents.
★ 323v2rayA. A web GUI client of Project V which supports VMess, VLESS, SS, SSR, Trojan, Tuic and Juicity protocols. 🚀
★ 15kXLeRobot. XLeRobot: Practical Dual-Arm Mobile Home Robot for $660
★ 5.4kSimplerEnv-OpenVLA. Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo, and OpenVLA) in simulation under common setups (e.g., Google Robot, WidowX+Bridge)
★ 271SpatialVLA. 🔥 SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.
★ 711cosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kFoundationStereo. [CVPR 2025 Best Paper Nomination] FoundationStereo: Zero-Shot Stereo Matching
★ 2.9kVision-Language-Models-Overview. A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
★ 683awesome-humanoid-robot-learning. A Paper List for Humanoid Robot Learning.
★ 2.6kPaper-List. A paper list of my history reading. Robotics, Learning, Vision.
★ 582Embodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15kRoboSpatial-Eval. Evaluation script for RoboSpatial-Home, a benchmark for spatial reasoning in 2D and 3D vision-language models.
★ 22open-eqa. OpenEQA Embodied Question Answering in the Era of Foundation Models
★ 365InternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kawesome-vlm-architectures. Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
★ 1.3kawesome-ai4s. AI for Science 论文解读合集(持续更新ing),论文/数据集/教程下载:hyper.ai
★ 3.3kai2thor. An open-source platform for Visual AI.
★ 1.8kFLaRe. [ICRA 25] FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
★ 49RL-VLM-F. Code for Reinforcement Learning from Vision Language Foundation Model Feedback
★ 140navchat. Code for ICRA24 paper "Think, Act, and Ask: Open-World Interactive Personalized Robot Navigation" Paper:https://arxiv.org/abs/2310.07968 Video:https://www.youtube.com/watch?v=rN5S8QIhhQc
★ 31LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kany4lerobot. 🎁 A collection of utilities for LeRobot.
★ 1.1kReasonFlux. [NeurIPS 2025 Spotlight] LLM post-training suite — featuring ReasonFlux, ReasonFlux-PRM, and ReasonFlux-Coder.
★ 541embodied-CoT. Embodied Chain of Thought: A robotic policy that reason to solve the task.
★ 410SimplerEnv. Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Google Robot, WidowX+Bridge) (CoRL 2024)
★ 1.1kopenvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7kopenpi. Python
★ 13kR1-V. Witness the aha moment of VLM with less than $3.
★ 4.1kGRAPE. GRAPE: Guided-Reinforced Vision-Language-Action Preference Optimization
★ 161AgiBot-World. [IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
★ 3.1kbypy. Python client for Baidu Yun (Personal Cloud Storage) 百度云/百度网盘Python客户端
★ 8.6kHOI4D-Instructions. The official instructions of HOI4D dataset.
★ 95RoboVLMs. Python
★ 474stable-diffusion. A latent text-to-image diffusion model
★ 73kdiffuser. Code for the paper "Planning with Diffusion for Flexible Behavior Synthesis"
★ 1.3kocto. Octo is a transformer-based robot policy trained on a diverse mix of 800k robot trajectories.
★ 1.7kDiffusionAsShader. [SIGGRAPH 2025] Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
★ 825ManiGaussian. [ECCV 2024] ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
★ 282FreeCloth. [CVPR 2025 Highlight] Official PyTorch implementation of "FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling"
★ 42VLM_survey. Collection of AWESOME vision-language models for vision tasks
★ 3.1kprismatic-vlms. A flexible and efficient codebase for training visually-conditioned language models (VLMs)
★ 1ksotopia-pi. Sotopia-π: Interactive Learning of Socially Intelligent Language Agents (ACL 2024)
★ 856d-object-pose-estimation. This repository summarizes papers and codes for 6D Object Pose Estimation.
★ 204DenseMatcher. Official Repository for DenseMatcher Learning 3D Semantic Correspondence for Category-Level Manipulation from One Demo
★ 123Autoregressive-Models-in-Vision-Survey. [TMLR 2025🔥] A survey for the autoregressive models in vision.
★ 805Human-Aware-Assistance-Codespace. Python
★ 7CleanDiffuser. CleanDiffuser: An Easy-to-use Modularized Library for Diffusion Models in Decision Making
★ 723DLInterview. Deep Learning Interview 深度学习面试题目汇总
★ 1.1kawesome-phd-advice. Collection of advice for prospective and current PhD students
★ 2.1kSAT-HMR. [CVPR 2025] SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens
★ 121co-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5kDINO. [ICLR 2023] Official implementation of the paper "DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection"
★ 2.8kdino. PyTorch code for Vision Transformers training with the Self-Supervised learning method DINO
★ 7.6kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kawesome-urdf. A curated list of Unified Robot Description Format (URDF) libraries, tools and resources.
★ 161point_odyssey. Official code for PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking (ICCV 2023)
★ 197act. Python
★ 2.1kdiff-gaussian-rasterization-w-depth. Cuda
★ 232robocook. Python
★ 94OP-Align. Python
★ 14TexPose. [CVPR 2023] Official repository for "TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation"
★ 28Social-CH. [NeurIPS 2023] PyTorch Implementation of "Social Motion Prediction with Cognitive Hierarchies"
★ 32YPPF. Yuanpei Profile
★ 39pymc. Bayesian Modeling and Probabilistic Programming in Python
★ 9.7k3D-Diffusion-Policy. [RSS 2024] 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
★ 1.4kGrounded-SAM-2. Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2
★ 3.7kElasticReconstruction. 3D reconstruction system to creating detailed scene geometry from range video.
★ 659Open3D-ML. An extension of Open3D to address 3D Machine Learning tasks
★ 2.3kdiffusion_policy. [RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
★ 4.4kpytorch_kinematics. Robot kinematics implemented in pytorch
★ 817mimicgen. This code corresponds to simulation environments used as part of the MimicGen project.
★ 626awe. Waypoint-Based Imitation Learning for Robotic Manipulation
★ 146DriveDownloader. Download resources from online storage with ONLY ONE command line!!
★ 149DaXBench. Python
★ 71softagent. Algorithms for deformable object manipulation benchmarked in SoftGym
★ 45SoftMAC. Code repository for our paper SoftMAC: Differentiable Soft Body Simulation with Forecast-based Contact Model and Two-way Coupling with Articulated Rigid Bodies and Clothes
★ 39softgym. SoftGym is a set of benchmark environments for deformable object manipulation.
★ 353VGPL-Dynamics-Prior. [ICML 2020] Visual Grounding of Learned Physical Models (Dynamics Prior)
★ 40robomimic. robomimic: A Modular Framework for Robot Learning from Demonstration
★ 1.5ksmalldiffusion. Simple and readable code for training and sampling from diffusion models
★ 777DPI-Net. [ICLR 2019] Learning Particle Dynamics for Manipulating Rigid Bodies, Deformable Objects, and Fluids
★ 140Holistic-Robot-Pose-Estimation. [ECCV 2024] PyTorch implementation of "Real-time Holistic Robot Pose Estimation with Unknown States"
★ 40PhysDreamer. Code for PhysDreamer
★ 630diffusion-literature-for-robotics. Summary of key papers and blogs about diffusion models to learn about the topic. Detailed list of all published diffusion robotics papers.
★ 920rh20t_api. Python
★ 109d3fields. [CoRL 24 Oral] D^3Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement
★ 185SpacetimeGaussians. [CVPR 2024] Spacetime Gaussian Feature Splatting for Real-Time Dynamic View Synthesis
★ 826Dynamic3DGaussians. Python
★ 2.3kgaussian-splatting. Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
★ 23k4DGaussians. [CVPR 2024] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
★ 3.9kPhysGaussian. [CVPR 2024 Highlight] PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
★ 1.4kawesome-3D-gaussian-splatting. Curated list of papers and resources focused on 3D Gaussian Splatting, intended to keep pace with the anticipated surge of research in the coming months.
★ 8.8kRoboEXP. [CoRL 2024] RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation
★ 132.tmux. Oh my tmux! My self-contained, pretty & versatile tmux configuration made with 💛🩷💙🖤❤️🤍
★ 25kSAMforMedPic. 北京大学2023年秋季学期 机器学习大作业
★ 2final-vcl. final project
★ 2calvin. CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
★ 964Awesome-Robotics-Foundation-Models.
★ 1.4kFoundationPose. [CVPR 2024 Highlight] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
★ 3.5kRoboGen. A generative and self-guided robotic agent that endlessly propose and master new skills.
★ 1.2kGNFactor. [CoRL 2023 Oral] GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
★ 140Awesome-Generalist-Robots-via-Foundation-Models. Paper list in the survey paper: Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
★ 468