This is your work, valued
RePHO. [CVPR 2026 Highlight] Repository for Recovering Physically Plausible Human-Object Interactions from Monocular Videos
★ 19PSDiffusion.
★ 18dingbang777. I’m an undergraduate at Shanghai Jiao Tong University. My interests lie in computer vision and deep generative models
★ 1dingbang. My personal repository
★ 1RePHO. [CVPR 2026 Highlight] Repository for Recovering Physically Plausible Human-Object Interactions from Monocular Videos
★ 19HumanEgo. HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
★ 462Motubrain. MotuBrain: An Advanced World Action Model for Robot Control
★ 49EgoX. Code for "EgoX: Egocentric Video Generation from a Single Exocentric Video"
★ 741DreamDojo. Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
★ 1klingbot-va. [RSS 2026] Causal video-action world model for generalist robot control
★ 1.7kgiga-world-policy. GigaWorld-Policy: An Efficient Action-Centered World–Action Model
★ 1.4krh20t_api. Python
★ 109FastWAM. Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
★ 1.2kdreamzero. Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
★ 2.5kricl_openpi. Python
★ 67icrt. [ICRA 2025] In-Context Imitation Learning via Next-Token Prediction
★ 120genie_sim. Simulation Platform from AgiBot
★ 1.3kresearch-skills. Claude Code skills for AI/ML researchers
★ 40SOMA-X. SOMA: Unifying Parametric Human Body Models
★ 713GENMO. Python
★ 488kairos. Official code for world model Kairos
★ 2.4kxfactor-nvs. Public code for XFactor: Introduces the first geometry-free model to achieve true self-supervised / pose-free Novel View Synthesis (NVS) by learning transferable latent camera pose representations.
★ 160cosmos-transfer1-diffusion-renderer. Cosmos-Transfer1-DiffusionRenderer: High-quality video de-lighting and re-lighting based on Cosmos video diffusion framework
★ 832HACO_RELEASE. [NeurIPS 2025] This repo is official PyTorch implementation of the paper "Learning Dense Hand Contact Estimation from Imbalanced Data".
★ 62holosoma. Python
★ 1.5kProtoMotions. ProtoMotions is a GPU-accelerated simulation and learning framework for training physically simulated digital humans and humanoid robots.
★ 2.2kAny6D. [CVPR 2025] Any6D: Model-free 6D Pose Estimation of Novel Objects
★ 478sam-3d-body. The repository provides code for running inference with the SAM 3D Body Model (3DB), links for downloading the trained model checkpoints and datasets, and example notebooks that show how to use the model.
★ 3.4kDepth-Anything-3. Depth Anything 3
★ 6kCONTHO_RELEASE. [CVPR 2024] This repo is official PyTorch implementation of Joint Reconstruction of 3D Human and Object via Contact-Based Refinement Transformer.
★ 112UniPhys. Python
★ 70PARC. PARC: Physics-based Augmentation with Reinforcement Learning for Character Controllers
★ 340diffusion-vas. [CVPR 2025] Official code for Using Diffusion Priors for Video Amodal Segmentation
★ 124vipe. ViPE: Video Pose Engine for Geometric 3D Perception
★ 2.1kPDP. Official code for PDP
★ 64UniEgoMotion. Code and data for UniEgoMotion (ICCV 2025)
★ 64Awesome-4D-Spatial-Intelligence. A curated list of awesome papers for reconstructing 4D spatial intelligence from video. (arXiv 2507.21045)
★ 515ScoreHOI. Official repository of ScoreHOI (ICCV 2025)
★ 16Awesome-Imitation-Learning. A curated list of awesome imitation learning resources and publications
★ 608InterPose. (3DV 2026) Pytorch implementation of “InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos”
★ 28SAM2-GUI. A GUI to interact with Meta's Segment Anything Model 2 (SAM2)
★ 5SkillMimic-V2. [SIGGRAPH 2025] SkillMimic-V2: Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations
★ 154InteractVLM. [CVPR 2025] InteractVLM: 3D Interaction Reasoning from 2D Foundational Models
★ 139gotrack. GoTrack: Generic 6DoF Object Pose Refinement and Tracking, CV4MR 2025
★ 101david. Official Repository for ICCV 2025 paper DAViD: Modeling Dynamic Affordance of 3D Objects using Pre-trained Video Diffusion Models
★ 85dinov3. Reference PyTorch implementation and models for DINOv3
★ 11kParaHome. Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions
★ 243PHC. Official Implementation of the ICCV 2023 paper: Perpetual Humanoid Control for Real-time Simulated Avatars
★ 1.3kFoundationPose. [CVPR 2024 Highlight] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
★ 3.5kGVHMR. Code for "GVHMR: World-Grounded Human Motion Recovery via Gravity-View Coordinates", Siggraph Asia 2024, TPAMI 2026
★ 1.8kIsaacGymEnvs. Isaac Gym Reinforcement Learning Environments
★ 3kviztracer. A debugging and profiling tool that can trace and visualize python code execution
★ 7.7kcosmos-transfer1. Cosmos-Transfer1 is a world-to-world transfer model designed to bridge the perceptual divide between simulated and real-world environments.
★ 812SMPLSim. Simulating SMPL humanoid, supporting PHC/PHC-MJX/PULSE/SimXR code bases.
★ 347InterAct. [CVPR 2025] InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
★ 198vggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kawesome-humanoid-robot-learning. A Paper List for Humanoid Robot Learning.
★ 2.6kAwesome-Controllable-Video-Generation. [ArXiv 2025] A survey about controllable video generation: This repo is the official awesome of "Controllable video generation: A survey"
★ 760hoiverse. Python
★ 3InterMimic. [CVPR 2025 Highlight] InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions
★ 520PSDiffusion.
★ 18open-3dhoi. Python
★ 31SkillMimic. [CVPR 2025 Highlight] SkillMimic: Learning Basketball Interaction Skills from Demonstrations
★ 430LearningHumans. Code snippets for understanding common techniques for virtual humans.
★ 130Matting-Anything. Matting Anything Model (MAM), an efficient and versatile framework for estimating the alpha matte of any instance in an image with flexible and interactive visual or linguistic user prompt guidance.
★ 715stable-diffusion-webui-Layer-Divider. Layer-Divider, an extension for stable-diffusion-webui using the segment-anything model (SAM)
★ 184