This is your work, valued
I am a graduate student from Tsinghua University. Interested in Robotics🤖. Intern @ BIGAI
AnyBimanual. [ICCV2025] AnyBimanual: Transfering Unimanual Policy for General Bimanual Manipulation
★ 102robotwin. Python
★ 5peract_bimanual. Python
★ 5Awesome-WholeBody-LocoManipulation.
★ 4diffad. Python
★ 2egoallo. Estimating Body and Hand Motion in an Ego-sensed World
★ 2lerobot_source. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 2glove_to_hand. Python
★ 2cobot_dp. Python
★ 1modelscope. Python
★ 1VLAS. forked from https://github.com/whichwhichgone/VLAS/tree/main
★ 1stablediffusion4ad. Python
★ 1openpi. Python
★ 1tau-0-wm. Python
★ 303DOMINO. [ECCV 2026] Towards Generalizable Robotic Manipulation in Dynamic Environments
★ 237SALMONN. SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5kPRISM. [IROS2026] PRISM: Precision and contact-rich Real-world Industrial Skill dataset with Multimodal sensing
★ 3LAPA. [ICLR 2025] LAPA: Latent Action Pretraining from Videos
★ 562glove_to_hand. Python
★ 2RynnBrain. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Models
★ 846Awesome-WAM. A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
★ 1.2kAwesome-WholeBody-LocoManipulation.
★ 4awesome-humanoid-robot-learning. A Paper List for Humanoid Robot Learning.
★ 2.6kProPainter. [ICCV 2023] ProPainter: Improving Propagation and Transformer for Video Inpainting
★ 6.8ksam3. The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 11kLIBERO-plus. Official repository of LIBERO-plus, a generalized benchmark for in-depth robustness analysis of vision-language-action models.
★ 400soma-retargeter. SOMA BVH to humanoid robot motion retargeting library built with Newton and NVIDIA Warp
★ 509CodePilot. A multi-model AI agent desktop client — connect any AI provider, extend with MCP & skills, control from your phone. Built with Electron + Next.js.
★ 6.3kacademic-research-skills. Academic Research Skills for Claude Code: research → write → review → revise → finalize
★ 40kGMR. [ICRA 2026] GMR: General Motion Retargeting. Retarget human motions into diverse humanoid robots in real time on CPU. Retargeter for TWIST.
★ 2.5kGR00T-WholeBodyControl. Welcome to GR00T Whole-Body Control (WBC)! This is a unified platform for developing and deploying advanced humanoid controllers. This includes: Decoupled WBC models used in NVIDIA Isaac-Gr00t, Gr00t N1.5 and N1.6 and GEAR-SONIC
★ 3kTWIST2. [arXiv 2025] TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
★ 859whole_body_tracking. Python
★ 2.3kDexUMI. DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation
★ 246hermes-agent. The agent that grows with you
★ 223kAuto-claude-code-research-in-sleep. ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
★ 14kcolleague-skill. 将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
★ 21kInternVLA-A-series. InternVLA-A1: Unifying Understanding, Generation, and Action for Robotic Manipulation
★ 522Psi0. [RSS26'] Welcome to Psi-Zero, a Humanoid VLA towards Universal Humanoid Intelligence.
★ 2.7kMotus. Official code of Motus: A Unified Latent Action World Model
★ 1.2kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kdreamzero. Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
★ 2.5klingbot-va. [RSS 2026] Causal video-action world model for generalist robot control
★ 1.7kcosmos-policy. Cosmos Policy
★ 8424DGaussians. [CVPR 2024] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
★ 3.9kvlash. Real-Time VLAs via Future-state-aware Asynchronous Inference.
★ 458DynamicVLA. The official implementation of "DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation". (arXiv 2601.22153)
★ 315lingbot-depth. Masked Depth Modeling for Spatial Perception
★ 1.5kPointWorld. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
★ 418cosmos-predict2.5. Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the form of video.
★ 1.3kWritingAIPaper. Writing AI Conference Papers: A Handbook for Beginners
★ 3.9kqc. Python
★ 397Awesome-RL-VLA. A Survey on Reinforcement Learning of Vision-Language-Action Models for Robotic Manipulation
★ 812Awesome-ML-SYS-Tutorial. My learning notes for ML SYS.
★ 6.8ksiiRL. siiRL: Shanghai Innovation Institute RL Framework for Advanced LLMs and Multi-Agent Systems
★ 368sam-3d-objects. SAM 3D Objects
★ 7.2kAwesome-Spatial-Intelligence-in-VLM. A paper list for spatial reasoning
★ 767MoMaGen. Python
★ 67BEHAVIOR-1K. BEHAVIOR-1K: a platform for accelerating Embodied AI research. Join our Discord for support: https://discord.gg/bccR5vGFEx
★ 1.6kOpenReal2Sim. A toolbox for real-to-sim reconstruction and robotic simulation
★ 237openpi. Python
★ 1diffad. Python
★ 2RoboOS. 🤖 RoboOS: A Universal Embodied Operating System for Cross-Embodied and Multi-Robot Collaboration
★ 605hello-agents. 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
★ 70kDINO-X-API. DINO-X: The World's Top-Performing Vision Model for Open-World Object Detection and Understanding
★ 1.4kmobipi. Official implementation for Mobi-π.
★ 128habitat-lab. A modular high-level library to train embodied AI agents across a variety of tasks and environments.
★ 3.1kDeformable-3D-Gaussians. [CVPR 2024] Official implementation of "Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction"
★ 1.2kShareRobot.
★ 62starVLA. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
★ 3.3k3D-R1. 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
★ 414RLVR-World. Official repository for "RLVR-World: Training World Models with Reinforcement Learning" (NeurIPS 2025), https://arxiv.org/abs/2505.13934
★ 272RLinf. RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
★ 4.4kDexGarmentLab. [NeurIPS 2025 Spotlight 🎊] DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy
★ 160GarmentLab. This is the official repository of GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation
★ 176MV-Adapter. [ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
★ 1.3kRoboFactory. [ICCV 2025] RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
★ 137VLA-R1. Python
★ 74verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kSimpleVLA-RL. [ICLR 2026] SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
★ 1.8kOmniVLA. Official repository for OmniVLA training and inference code
★ 310VLA-Adapter. VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
★ 2.3kslime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kdppo. Official implementation of Diffusion Policy Policy Optimization, arxiv 2024
★ 842RDT2. Official code of RDT 2
★ 798RynnVLA-001. [ICRA 2026] RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
★ 303unifolm-world-model-action. Python
★ 1.1kphantom. Python
★ 109Depth-Anything-V2. [NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
★ 8.6kE2FGVI. Official code for "Towards An End-to-End Framework for Flow-Guided Video Inpainting" (CVPR2022)
★ 1.2kdetectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35kTRELLIS. Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
★ 13kOnePoseviaGen. [CORL 2025 Oral]One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation.
★ 486modelscope. Python
★ 1Grounded-SAM-2. Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2
★ 3.7kRoboBrain2.5. RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. 🎉🎉🎉
★ 1.1kLIBERO. Benchmarking Knowledge Transfer in Lifelong Robot Learning
★ 2.1kGenPose2. [ECCV 2024] GenPose++: A generative category-level 6D object pose estimation and tracking approach proposed in Omni6DPose.
★ 111Inpaint-Anything. Inpaint anything using Segment Anything and inpainting models.
★ 7.7klama. 🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022
★ 10kdexmv-sim. DexMV: Imitation Learning for Dexterous Manipulation from Human Videos, ECCV 2022
★ 198rigvid. Python
★ 62rovi-aug. Augment robotics demonstration datasets with different robots and viewpoints
★ 42mirage. Mirage: a zero-shot cross-embodiment policy transfer method. Benchmarking code for cross-embodiment policy transfer.
★ 34nora. NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
★ 222EgoHOS. Fine-Grained Egocentric Hand-Object Segmentation, ECCV 2022
★ 144embodied_reasoner. Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
★ 205hume. 🦾 A Dual-System VLA with System2 Thinking
★ 148VLAS. forked from https://github.com/whichwhichgone/VLAS/tree/main
★ 1TinyLLaVA-Video-R1. TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
★ 116ManiGaussian_Bimanual. [IROS 2025] ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
★ 46pytorch3d. PyTorch3D is FAIR's library of reusable components for deep learning with 3D data
★ 9.9kRoboTwin. [ICML 2026] RoboTwin 2.0 Offical Repo
★ 2.7kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kespnet. End-to-End Speech Processing Toolkit
★ 9.9kcalvin. CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
★ 964LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25ktextgrad. TextGrad: Automatic ''Differentiation'' via Text -- using large language models to backpropagate textual gradients. Published in Nature.
★ 3.7kVoyager. An Open-Ended Embodied Agent with Large Language Models
★ 7.1kVLAS. Python
★ 48vjepa2. PyTorch code and models for VJEPA2 self-supervised learning from video.
★ 4.4kmetamorph. Code for "MetaMorph: Learning Universal Controllers with Transformers", Gupta et al, ICLR 2022
★ 131egoallo. Estimating Body and Hand Motion in an Ego-sensed World
★ 2ManiSkill. SAPIEN Manipulation Skill Framework, an open source GPU parallelized robotics simulator and benchmark, led by Hillbot, Inc.
★ 2UniVLA. [RSS 2025] Learning to Act Anywhere with Task-centric Latent Actions
★ 1.1kMotionBERT. [ICCV 2023] PyTorch Implementation of "MotionBERT: A Unified Perspective on Learning Human Motion Representations"
★ 1.4krobocasa. RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
★ 1.6kMiniCPM. MiniCPM5-1B: A SOTA 1B on-device LLM, small yet powerful.
★ 10krobotwin. Python
★ 5OneTwoVLA. Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"
★ 236UniSkill. [CoRL 2025] UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
★ 88LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kreal2render2real. [CoRL 2025] Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
★ 367FoundationPose. [CVPR 2024 Highlight] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
★ 3.5kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kIsaac-GR00T. NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7.7kFoundationStereo. [CVPR 2025 Best Paper Nomination] FoundationStereo: Zero-Shot Stereo Matching
★ 2.9khamer. HaMeR: Reconstructing Hands in 3D with Transformers
★ 1.1khumor. Code for ICCV 2021 paper "HuMoR: 3D Human Motion Model for Robust Pose Estimation"
★ 577openvla-oft. Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
★ 1.3kegoallo. Estimating Body and Hand Motion in an Ego-sensed World
★ 289segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kopen-pi-zero. Re-implementation of pi0 vision-language-action (VLA) model from Physical Intelligence
★ 1.5kegoego_release. Official Implementation of the Paper: Ego-Body Pose Estimation via Ego-Head Pose Estimation (CVPR 2023 Award Candidate)
★ 148vq_bet_official. Official code for "Behavior Generation with Latent Actions" (ICML 2024 Spotlight)
★ 211ManipTrans. [CVPR 2025] 🎉 Official repository of "ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning"
★ 330Qwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kTesserAct. ICCV 2025 | TesserAct: Learning 4D Embodied World Models
★ 404DexGraspVLA. [AAAI'26 Oral] DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
★ 556lerobot_source. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 2anygrasp_sdk. Python
★ 953EgoH4. Official code releasse for "The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation"
★ 33fakemailaddress. Python
★ 2RoboSplat. [RSS 2025] Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation
★ 198Dummy-Robot. 我的超迷你机械臂机器人项目。
★ 15klerobot_piper. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 87UniHSI. [ICLR 2024 Spotlight] Unified Human-Scene Interaction via Prompted Chain-of-Contacts
★ 246OpenHomie. Open-sourced code for "HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit".
★ 596mimicgen. This code corresponds to simulation environments used as part of the MimicGen project.
★ 626dexmimicgen. This code corresponds to simulation environments used as part of the DexMimicGen project.
★ 266TeleVision. [CoRL 2024] Open-TeleVision: Teleoperation with Immersive Active Visual Feedback
★ 1.3kOpen-Teach. A Versatile Teleoperation framework for Robotic Manipulation using Meta Quest3
★ 417ASAP. [RSS 2025] "ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills"
★ 2.1kunitree_lerobot. The unitree_il_lerobot open-source project is a modification of the LeRobot open-source training framework, enabling the training and testing of data collected using the dual-arm dexterous hands of Unitree's G1 robot.
★ 729roboengine. Official Reporsitory of "RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation"
★ 165Depth-Anything. [CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation
★ 8.2kHSMR. [CVPR25 Oral (Top 3.3%)] Official code for paper "Reconstructing Humans with a Biomechanically Accurate Skeleton".
★ 632vlarl. Single-file implementation to advance vision-language-action (VLA) models with reinforcement learning.
★ 447sam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kMiniCPM-V. A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
★ 26krobotics_arXiv_daily. Fetching Embodied AI Paper from ArXiv automatically
★ 331Awesome-RL-for-LRMs. A Survey of Reinforcement Learning for Large Reasoning Models
★ 2.5khuman-policy. Python
★ 257smplify-x. Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
★ 2.2kboss. Code for the paper Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance, accepted to CoRL 2023 as an Oral Presentation.
★ 35lemon_3d. Python
★ 94posenet-pytorch. A PyTorch port of Google TensorFlow.js PoseNet (Real-time Human Pose Estimation)
★ 310RNNPose. RNNPose: Recurrent 6-DoF Object Pose Refinement with Robust Correspondence Field Estimation and Pose Optimization, CVPR 2022
★ 189yt-dlp. A feature-rich command-line audio/video downloader
★ 182kagora_evaluation. Python
★ 117smplx. SMPL-X
★ 2.7kDeep3D. Real-Time end-to-end 2D-to-3D Video Conversion, based on deep learning.
★ 231deep3d. Automatic 2D-to-3D Video Conversion with CNNs
★ 1.3kopen_x_embodiment. Jupyter Notebook
★ 2kEgo4d. Ego4d dataset repository. Download the dataset, visualize, extract features & example usage of the dataset
★ 627ReconX. [TIP 2026] ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
★ 709awesome-embodied-vla-va-vln. A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
★ 3.4kguided-diffusion. Python
★ 7.4kgraspnetAPI. Toolbox for our GraspNet-1Billion dataset.
★ 361SoM. [arXiv 2023] Set-of-Mark Prompting for GPT-4V and LMMs
★ 1.6kvisualnav-transformer. Official code and checkpoint release for mobile robot foundation models: GNM, ViNT, and NoMaD.
★ 1.3kHybrid-VLA. HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
★ 352LIFT3D. [CVPR 2025]Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
★ 185mediapy. This Python library makes it easy to display images and videos in a notebook.
★ 450dummy_ctrl. Python
★ 2Dita. ICCV2025
★ 171Customize-it-3D. [ICRA 2025] Official implementation of Customize-It-3D: High-Quality 3D Creation from A Single Image Using Subject-Specific Knowledge Prior
★ 75universal_manipulation_interface. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
★ 1.5kEXPRESS-Bench. Embodied Question Answering (EQA) benchmark and method (ICCV 2025)
★ 60vscodium. binary releases of VS Code without MS branding/telemetry/licensing
★ 33kBimanual_ur5e_joystick_control. Python
★ 5robomimic. robomimic: A Modular Framework for Robot Learning from Demonstration
★ 1.5kact. Python
★ 2.1kmujoco_menagerie. A collection of high-quality models for the MuJoCo physics engine, curated by Google DeepMind.
★ 3.8kmujoco. Multi-Joint dynamics with Contact. A general purpose physics simulator.
★ 14kCogACT. A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
★ 429diffusion_policy. [RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
★ 4.4kgoogle-research. Google Research
★ 38knomad. Nomad is an easy-to-use, flexible, and performant workload orchestrator that can deploy a mix of microservice, batch, containerized, and non-containerized applications. Nomad is easy to operate and scale and has native Consul and Vault integrations.
★ 17klerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kros2. The Robot Operating System, is a meta operating system for robots.
★ 5.8kRoboFlamingo. Code for RoboFlamingo
★ 437act-plus-plus. Imitation learning algorithms with Co-training for Mobile ALOHA: ACT, Diffusion Policy, VINN
★ 3.6k