This is your work, valued
Transformer-in-Vision. Recent Transformer-based CV and related works.
★ 1.3kLLM-in-Vision. Recent LLM-based CV and related works. Welcome to comment/contribute!
★ 871HOI-Learning-List. A list of Human-Object Interaction Learning.
★ 714Transferable-Interactiveness-Network. Code for Transferable Interactiveness Knowledge for Human-Object Interaction Detection. (CVPR'19, TPAMI'21)
★ 239HAKE-Action-Torch. HAKE-Action in PyTorch
★ 238HAKE. HAKE: Human Activity Knowledge Engine (CVPR'18/19/20, NeurIPS'20, TPAMI'21)
★ 235DJ-RN. As a part of HAKE project (HAKE-3D). Code for our CVPR2020 paper "Detailed 2D-3D Joint Representation for Human-Object Interaction".
★ 103HAKE-Action. As a part of the HAKE project, includes the reproduced SOTA models and the corresponding HAKE-enhanced versions (CVPR2020).
★ 100SymNet. As a part of the HAKE project (HAKE-Object), code for SymNet (CVPR'20 and TPAMI'21).
★ 53Sandwich. Bidirectional Mapping between Action Physical-Semantic Space
★ 34HAKE-AVA. Shell
★ 31sampling-argmax. Code for "Localization with Sampling-Argmax", NeurIPS 2021
★ 10SRDA-ECCV2018. ECCV2018: SRDA: Generating Instance Segmentation Annotation via Scanning, Reasoning and Domain Adaptation
★ 8labelKeypoint. Keypoints label tools
★ 3Mask_RCNN. Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow(tf1.3.0+cuda8+cudnn6)
★ 2BBox-Label-Tool. A simple tool for labeling object bounding boxes in images
★ 2AlphaPose. Multi-Person Pose Estimation System
★ 2HAT-4D-Lifting-Monocular-Video-for-4D-Multi-Object-Interactions-via-Human-Agent-Collaboration. Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
★ 7GR00T-WholeBodyControl. Welcome to GR00T Whole-Body Control (WBC)! This is a unified platform for developing and deploying advanced humanoid controllers. This includes: Decoupled WBC models used in NVIDIA Isaac-Gr00t, Gr00t N1.5 and N1.6 and GEAR-SONIC
★ 3kChronoFlow-Policy. Official implementation of ChronoFlow-Policy: a diffusion-based visuomotor policy that jointly models past-current-future object-gripper interaction flows for robot manipulation.
★ 6Wh0. Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data
★ 18Geometric_Primary_Structure. Code for CVPR2026 Paper "Revisiting Articulated Parts Perception in Robot Manipulation"
★ 8expo-ft. Python
★ 73Open-d4rt. Python
★ 838ViFailback. [CVPR 2026] Official repository for the paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols".
★ 16OmniXtreme. Python
★ 714lenny-skills. 86 product management skills from Lenny's Podcast for Claude Code and AI agents. Hiring, user research, strategy, shipping, and more.
★ 1.2kGM-X. Official Repo of The Great March Project. https://www.rhos.ai/research/gm-100
★ 22lingbot-vla. A Pragmatic VLA Foundation Model
★ 1.7kIPR-1. Official Repo of IPR-1 Project. https://www.rhos.ai/research/ipr-1
★ 30garmagenet-impl. Official implementation for GarmageNet: A Multimodal Generative Framework for Sewing Pattern Design and Generic Garment Modeling (SIGGRAPH Asia 2025)
★ 654dhoi_autorecon. Python
★ 7RHOS. RHOS Lab
★ 5open4dhoi_code. Python
★ 50exUMI. exUMI (CoRL 2025)
★ 71RoboHiMan. RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
★ 17ORoboSoul.
★ 16MBA. [RA-L 2025 & ICRA 2026] :kissing_cat: Motion Before Action: Diffusing Object Motion as Manipulation Condition
★ 73SIME. [IROS 2025] SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
★ 17SmartPlay. SmartPlay is a benchmark for Large Language Models (LLMs). Uses a variety of games to test various important LLM capabilities as agents. SmartPlay is designed to be easy to use, and to support future development of LLMs.
★ 146TactAR_APP. [RSS 2025] TactAR teleopeartion APP in "Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation"
★ 86reactive_diffusion_policy. [RSS 2025] Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
★ 365Awesome-Embodied-AI-Job. Lumina Robotics Talent Call | Lumina社区具身智能招贤榜 | A list for Embodied AI / Robotics Jobs (PhD, RA, intern, etc
★ 1.5kopen-3dhoi. Python
★ 31DensePolicy2D. [ICCV 2025] :bouquet: 2D version of Dense Policy (DSP)
★ 34DensePolicy. [ICCV 2025] :bouquet: Dense Policy (DSP): Bidirectional Autoregressive Learning of Actions
★ 79M3VOS_Experiment. Jupyter Notebook
★ 10SemiAuto-Multi-Level-Annotation-Tool. This repo is the office implement of a Semi-Auto Multi-Level Annotation Tool used in M-cube-VOS. It allows for user to annotate the mask of target objects efficiently.
★ 4mujoco_playground. An open-source library for GPU-accelerated robot learning and sim-to-real transfer.
★ 2.1kNSFC-LaTex. BibTeX Style
★ 1.6kopenpi. Python
★ 13kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kOpen-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kOpen-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kAwesome-Human-Motion. An aggregation of human motion understanding research.
★ 287VidTok. a family of versatile and state-of-the-art video tokenizers.
★ 454diamond. DIAMOND (DIffusion As a Model Of eNvironment Dreams) is a reinforcement learning agent trained in a diffusion world model. NeurIPS 2024 Spotlight.
★ 2.1kGenieRedux. A framework for training world models with virtual environments, complete with annotated environment dataset (RetroAct), exploration agent (AutoExplore Agent), and GenieRedux-G - an implementation of Genie with enhancements
★ 77ADAM. We introduce ADAM, An emboDied causal Agent in Minecraft, that can autonomously navigate the open world, perceive multimodal contexts, learn causal world knowledge, and tackle complex tasks through lifelong learning.
★ 33Playable-Game-Generation. An open-source lightweight game generation paradigm. It includes everything from data processing to model architecture design and playability-based evaluation methods. The game runs at 20 FPS on a single consumer-grade graphics card (RTX-2060) while maintaining high playability.
★ 120HunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kVisionZip. Official repository for VisionZip (CVPR 2025)
★ 442HDyS. Homogeneous Dynamics Space for Heterogeneous Humans (CVPR 2025)
★ 9ImDy. ImDy: Human Inverse Dynamics from Imitated Observations (ICLR 2025)
★ 20LLM_Inception. [ICLR 2025] This repo is the official implementation of "The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs".
★ 13TwoStageReason. Official implementation of ECCV 2024 paper: Take A Step Back: Rethinking the Two Stages in Visual Reasoning
★ 13lerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kMotion-Occupancy-Base. Official repository for gathering data of Revisit Human-Scene Interaction via Space Occupancy (ECCV 2024).
★ 30HumanVLA. Python
★ 135nsfc. nsfc - 国家自然科学基金项目LaTeX模版(面地青CBA)
★ 1.3kSymbol-LLM. Code for NeurIPS2023 Paper "Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity Reasoning"
★ 26LoRS_Distill. Code for our ICML'24 on multimodal dataset distillation
★ 44EgoPCA. This is official code of ICCV'23 paper: EgoPCA: A New Framework for Egocentric Hand-Object Interaction Understanding.
★ 2video_distillation. Official implementation of Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement.
★ 32Grounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kInternLM-XComposer. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2.9kMultiPLY. Code for MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
★ 135myosuite. MyoSuite is a collection of environments/tasks to be solved by musculoskeletal models simulated with the MuJoCo physics engine and wrapped in the OpenAI gym API.
★ 1.2ksearch_with_lepton. Building a quick conversation-based search demo with Lepton AI.
★ 8.1kWonder3D. Single Image to 3D using Cross-Domain Diffusion for 3D Generation
★ 5.4kocto. Octo is a transformer-based robot policy trained on a diverse mix of 800k robot trajectories.
★ 1.7kObjectConceptLearning. This the official repository of OCL (ICCV 2023).
★ 25GoldFromOres-BiLP. Preview code of ECCV'24 paper "Distill Gold from Massive Ores" (BiLP)
★ 25awesome-video-self-supervised-learning. A curated list of awesome self-supervised learning methods in videos
★ 172DiffHOI. Official implementation of the paper "Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model"
★ 67Pose_to_SMPL. A tool to fit SMPL parameters from 3D-pose datasets that contain key-points of human body.
★ 155Total-Recon. [ICCV 2023] Total-Recon: Deformable Scene Reconstruction for Embodied View Synthesis
★ 216taichi. Productive, portable, and performant GPU programming in Python.
★ 28klerf. Code for LERF: Language Embedded Radiance Fields
★ 731Anything-3D. Segment-Anything + 3D. Let's lift anything to 3D.
★ 1.6kstable-dreamfusion. Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
★ 8.8ktaichi-nerfs. Implementations of NeRF variants based on Taichi + PyTorch
★ 828Segment-Everything-Everywhere-All-At-Once. [NeurIPS 2023] Official implementation of the paper "Segment Everything Everywhere All at Once"
★ 4.8kEditAnything. Edit anything in images powered by segment-anything, ControlNet, StableDiffusion, etc. (ACM MM)
★ 3.4ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kVL-CheckList. Evaluating Vision & Language Pretraining Models with Objects, Attributes and Relations. [EMNLP 2022]
★ 138open_flamingo. An open-source framework for training large multimodal models.
★ 4.1kLLM-in-Vision. Recent LLM-based CV and related works. Welcome to comment/contribute!
★ 871AlphaTracker. AlphaTracker is a computer vision pipeline with the practical and real-time advantages , which requires minimal hardware requirements and produces reliable tracking of multiple unmarked animals. An easy-to-use user interface further enables manual inspection and curation of results.
★ 75Transformer-in-Vision. Recent Transformer-based CV and related works.
★ 1.3kDINO. [ICLR 2023] Official implementation of the paper "DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection"
★ 2.8kdetrex. detrex is a research platform for DETR-based object detection, segmentation, pose estimation and other visual recognition tasks.
★ 2.3kmvig-rhos.github.io. HTML
★ 6Robotics-Conference-Acceptance-Rate. Statistics of acceptance rate for the main robotics conference
★ 9DnD. Code for "D&D: Learning Human Dynamics from Dynamic Camera", ECCV 2022 Oral
★ 117aitviewer. A set of tools to visualize and interact with sequences of 3D data.
★ 744unity-antagonistic-controller. Generating Upper-Body Motion for Real-Time Characters Making their Way through Dynamic Environments - 2022 - SCA
★ 124YLearn. YLearn, a pun of "learn why", is a python package for causal inference
★ 433GAP. official implementation for Language Supervised Training for Skeleton-based Action Recognition
★ 131Sandwich. Bidirectional Mapping between Action Physical-Semantic Space
★ 34alpa. Training and serving large-scale neural networks with auto parallelization.
★ 3.2kacademia-hugo. Academia is a Hugo resume theme. You can showcase your academic resume, publications and talks using this theme.
★ 229Awesome-Dataset-Distillation. A curated list of awesome papers on dataset distillation and related applications.
★ 2kGLAMR. [CVPR 2022 Oral] Official PyTorch Implementation of "GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras”.
★ 389Partial_Distance_Correlation. This is the official GitHub for paper: On the Versatile Uses of Partial Distance Correlation in Deep Learning, in ECCV 2022
★ 176ROMP. Monocular, One-stage, Regression of Multiple 3D People and their 3D positions & trajectories in camera & global coordinates. ROMP[ICCV21], BEV[CVPR22], TRACE[CVPR2023]
★ 1.5kDLSA. The code accompanying our ECCV'22 papers: Constructing Balance from Imbalance for Long-tailed Image Recognition
★ 18Body-Part-Map-for-Interactiveness. Code for ECCV2022 Paper "Mining Cross-Person Cues for Body-Part Interactiveness Learning in HOI Detection"
★ 37parti.
★ 1.6kInteractiveness-Field. Python
★ 7ZeroVL. [ECCV2022] Contrastive Vision-Language Pre-training with Limited Resources
★ 46DCR. Official implementation of our CVPR'22 paper.
★ 13open_clip. An open source implementation of CLIP.
★ 14kmtt-distillation. Official code for our CVPR '22 paper "Dataset Distillation by Matching Training Trajectories"
★ 441CPPF. CPPF: Towards Robust Category-Level 9D Pose Estimation in the Wild (CVPR2022)
★ 57OC-Immunity. Python
★ 8HAKE-AVA. Shell
★ 31sampling-argmax. Code for "Localization with Sampling-Argmax", NeurIPS 2021
★ 10ICON. [CVPR'22] ICON: Implicit Clothed humans Obtained from Normals
★ 1.7kawesome-3dbody-papers. 😎Awesome list of papers about 3D body
★ 662sampling-argmax. Code for "Localization with Sampling-Argmax", NeurIPS 2021
★ 92pysot. SenseTime Research platform for single object tracking, implementing algorithms like SiamRPN and SiamMask.
★ 1hellosiyuan. Python
★ 3awesome-long-tail-learning.
★ 490model-vs-human. Benchmark your model on out-of-distribution datasets with carefully collected human comparison data (NeurIPS 2021 Oral)
★ 362AI4Animation. Bringing Characters to Life with Computer Brains in Unity
★ 8.8khotr. Official repository for HOTR: End-to-End Human-Object Interaction Detection with Transformers (CVPR'21, Oral Presentation)
★ 154SlowFast. PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
★ 7.4kConsNet. 🚴♂️ ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction Detection (MM 2020)
★ 35ahp. AHP: Amodal Human Perception dataset
★ 15PGT. Official code of paper "PGT: A Progressive Method for Training Models on Long Videos" on CVPR2021
★ 30awesome-holistic-3d. A list of papers and resources (data,code,etc) for holistic 3D reconstruction in computer vision
★ 646Handeye-Calibration-ROS. 🤖 Elaborated hand-eye calibration tutorials (ROS-binding)
★ 125MySummary. My Design Philosophy Summary (Most of them are in Chinese)
★ 537CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kgithub-buttons. Showcase the success of any GitHub repo or user with these simple, static buttons with dynamic counts.
★ 2.9kCanonicalVoting. Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes (CVPR2022)
★ 49UKPGAN. UKPGAN: A General Self-Supervised Keypoint Detector (CVPR2022)
★ 74OneNet. [ICML2021] What Makes for End-to-End Object Detection
★ 641SJTUThesis. 上海交通大学 LaTeX 论文模板 | Shanghai Jiao Tong University LaTeX Thesis Template
★ 3.8kHAKE-Action-Torch. HAKE-Action in PyTorch
★ 238object_level_visual_reasoning. Pytorch Implementation of "Object level Visual Reasoning in Videos", F. Baradel, N. Neverova, C. Wolf, J. Mille, G. Mori , ECCV 2018
★ 169Halpe-FullBody. Halpe: full body human pose estimation and human-object interaction detection dataset
★ 371human_body_prior. VPoser: Variational Human Pose Prior
★ 981zero_shot_hoi. Discovering human interaction with novel objects via zero-shot learning, CVPR, 2020
★ 42AlphAction. Spatio-Temporal Action Localization System
★ 427taskmodularnets. A set of neural network modules, which are small fully connected layers operating in semantic concept space. These modules are configured through a gating function conditioned on the task to produce features representing the compatibility between the input image and the concept under consideration.
★ 61HOI-Learning-List. A list of Human-Object Interaction Learning.
★ 714SymNet. As a part of the HAKE project (HAKE-Object), code for SymNet (CVPR'20 and TPAMI'21).
★ 53DJ-RN. As a part of HAKE project (HAKE-3D). Code for our CVPR2020 paper "Detailed 2D-3D Joint Representation for Human-Object Interaction".
★ 103HAKE-Action. As a part of the HAKE project, includes the reproduced SOTA models and the corresponding HAKE-enhanced versions (CVPR2020).
★ 100Few-Shot-Object-Detection-Dataset.
★ 394attributes-as-operators. Attribute-Object Visual Composition using Attributes as Operators
★ 65furniture. IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks
★ 566awesome-point-cloud-analysis. A list of papers and datasets about point cloud analysis (processing)
★ 4.2ksmplify-x. Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
★ 2.2kInstaboost. Code for ICCV2019 paper "InstaBoost: Boosting Instance Segmentation Via Probability Map Guided Copy-Pasting"
★ 402HAKE. HAKE: Human Activity Knowledge Engine (CVPR'18/19/20, NeurIPS'20, TPAMI'21)
★ 235labelKeypoint. Keypoints label tools
★ 3BBox-Label-Tool. A simple tool for labeling object bounding boxes in images
★ 2AlphaPose. Multi-Person Pose Estimation System
★ 2Mask_RCNN. Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow(tf1.3.0+cuda8+cudnn6)
★ 2video-long-term-feature-banks. Long-Term Feature Banks for Detailed Video Understanding
★ 383SRDA-ECCV2018. ECCV2018: SRDA: Generating Instance Segmentation Annotation via Scanning, Reasoning and Domain Adaptation
★ 8Transferable-Interactiveness-Network. Code for Transferable Interactiveness Knowledge for Human-Object Interaction Detection. (CVPR'19, TPAMI'21)
★ 239GAN_Theories. Resources and Implementations of Generative Adversarial Nets which are focusing on how to stabilize training process and generate high quality images: DCGAN, WGAN, EBGAN, BEGAN, etc.
★ 191iCAN. [BMVC 2018] iCAN: Instance-Centric Attention Network for Human-Object Interaction Detection
★ 264PRIN. Pointwise Rotation-Invariant Network (AAAI 2020)
★ 88PRNet. Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network (ECCV 2018)
★ 5kjsonlab. JSONLab: compact, portable, robust JSON/binary-JSON encoder/decoder for MATLAB/Octave
★ 320labelKeypoint. Keypoints label tools
★ 89bundler_sfm. Bundler Structure from Motion Toolkit
★ 1.6kai-deadlines. :alarm_clock: AI conference deadline countdowns
★ 6kAlphaPose. Real-Time and Accurate Full-Body Multi-Person Pose Estimation&Tracking System
★ 8.6kthe-gan-zoo. A list of all named GANs!
★ 15kco-fusion. Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects
★ 516MNC. Instance-aware Semantic Segmentation via Multi-task Network Cascades
★ 494openMVS. open Multi-View Stereo reconstruction library
★ 4.1kvaeblog.
★ 78fast-neural-style. Feedforward style transfer
★ 4.4kPhotographicImageSynthesis. Photographic Image Synthesis with Cascaded Refinement Networks
★ 1.2kopencv. Open Source Computer Vision Library
★ 90kcaffe. Caffe: a fast open framework for deep learning.
★ 35k