This is your work, valued
Ph.D. @ UMassο½B.Eng. @ SJTUο½Visiting Student @ MIT
PyBlend. PyBlend: a package for Blender with Python π¨
β 133SJTU-Course-Stack. SJTU-Course-Stack
β 8EE367. Stanford EE367 / CS448I: Computational Imaging
β 6GEFF. π GEFF: Gaze Estimation with Fused Features. Also the Project of AI2611
β 3AI2613-Homework. SJTU AI2613 Stochastic Processes
β 2Self-Learning. π€ SLS: Self-Learning Stack
β 1Vim-LaTeX-Configuration. Vim Script
β 1genai_workshop.github.io. JavaScript
β 1anyeZHY.github.io. SCSS
β 1SimpleVLA-RL. [ICLR 2026] SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
β 1.8kActionImages. Python
β 71vggt-omega. [CVPR 2026 Oral] VGGT Omega
β 3.8kverl-omni. Multimodal RL training framework for diffusion & omni models
β 691flash-attention-prebuild-wheels. Provide with pre-build flash-attention 2 and 3 package wheels on Linux and Windows using GitHub Actions
β 1.7kFD-Loss. Python
β 549UniDex. [CVPR 2026] UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos
β 169LoGeR. Reimplementation of LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
β 609WorldArena. WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
β 251Genesis-Humanoid. An all-in-one humanoid research platform on top of Genesis.
β 156cosmos-policy. Cosmos Policy
β 842VLM4VLA. Implementation of VLM4VLA
β 165TurboDiffusion. TurboDiffusion: 100β200Γ Acceleration for Video Diffusion Models
β 3.6kWan-Move. [NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
β 649BasicTS. A Fair and Scalable Time Series Forecasting Benchmark and Toolkit.
β 1.8kVideo-As-Prompt. [ICLR 2026] Official repo for paper "Video-As-Prompt: Unified Semantic Control for Video Generation"
β 443Awesome-World-Models. A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
β 3.3kTraceAnything. [ICLR 2026] Trace Anything: Representing Any Video in 4D via Trajectory Fields
β 543Awesome-Embodied-World-Model. Awesome paper list and repos of the paper "A comprehensive survey of embodied world models".
β 131openpi. Python
β 13kIR3D-Bench. [NeurIPS DB 2025] IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
β 464DNeX. 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
β 840dinov3. Reference PyTorch implementation and models for DINOv3
β 11kHumanRobotAlign. This is the official repo for [CVPR 2025] paper, Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation. https://jiaming-zhou.github.io/projects/HumanRobotAlign/
β 31vipe. ViPE: Video Pose Engine for Geometric 3D Perception
β 2.1kEWMBench. Official code for EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
β 129Genie-Envisioner-V1. Python
β 565gpt-oss. gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
β 20kDensePolicy. [ICCV 2025] :bouquet: Dense Policy (DSP): Bidirectional Autoregressive Learning of Actions
β 79viser. Web-based 3D visualization in Python
β 2.7kWan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
β 17kEasyCache. Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
β 292Awesome-World-Model. Collect some World Models for Autonomous Driving (and Robotic, etc.) papers.
β 2.2kscene-language. (CVPR 2025 Highlight) The Scene Language: Representing Scenes with Programs, Words, and Embeddings
β 264Pi3. [ICLR 2026] Ο^3: Permutation-Equivariant Visual Geometry Learning
β 2.1kMindJourney. [NeurIPS 2025] Source codes for the paper "MindJourney: Test-Time Scaling with World Models for Spatial Reasoning"
β 151DiffSynth-Studio. Enjoy the magic of Diffusion models!
β 13kPerspectiveFields. [CVPR 2023 Highlight] Perspective Fields for Single Image Camera Calibration
β 314MoGe. [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
β 2.7k4dgen. [ICLR 2026] Codebase for paper "Geometry-aware 4D Video Generation for Robot Manipulation"
β 1234DGC. Official implement of 4DGC(CVPR2025)
β 29Awesome-World-Models. A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related websites.
β 1.9kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
β 20kTaylorSeer. [ICCV2025] From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
β 409TeaCache. Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
β 1.4kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
β 17kATI. Official implementation of ATI: Any Trajectory Instruction for Controllable Video Generation. https://arxiv.org/pdf/2505.22944
β 356OctoThinker. Revisiting Mid-training in the Era of Reinforcement Learning Scaling
β 189RynnVLA-002. RynnVLA-002: A Unified Vision-Language-Action and World Model
β 1.1kDetAny3D. [ICCV 2025] Detect Anything 3D in the Wild
β 288gemini-robotics-sdk. Python
β 590diffusion-forcing. code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"
β 1.3kfast-spatial-mem. Fast Spatial Memory with Elastic Test-Time Training (4D-LRM + 4D-LVSM)
β 104RoboTwin. [ICML 2026] RoboTwin 2.0 Offical Repo
β 2.7kEmbodied-Web-Agent. Ruby
β 40Virtual-Community. Virtual Community: An Open World for Humans, Robots, and Society
β 192Mirage. [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
β 294GR00T-Dreams. DreamGen: Nvidia GEAR Lab's initiative to solve the robotics data problem using world models
β 593CommVQ. [ICML 2025] CommVQ: Commutative Vector Quantization for KV Cache Compression
β 283dbinpacking. A python library for 3D Bin Packing
β 457PartPacker. Efficient Part-level 3D Object Generation via Dual Volume Packing
β 824cosmos-predict2. Cosmos-Predict2 is a collection of general-purpose world foundation models for Physical AI that can be fine-tuned into customized world models for downstream applications.
β 7933D-bin-packing. 3D Bin Packing improvements based on https://github.com/enzoruiz/3dbinpacking
β 288WiLoR. WiLoR: End-to-end 3D hand localization and reconstruction in-the-wild
β 614PVSGAnnotation. Python
β 73DFlowAction. Python
β 62AdaWorld. [ICML'25] The PyTorch implementation of paper: "AdaWorld: Learning Adaptable World Models with Latent Actions".
β 254DeepVerse. DeepVerse: 4D Autoregressive Video Generation as a World Model
β 230ORV. [CVPR 2026] ORV: 4D Occupancy-centric Robot Video Generation.
β 108LaCT. Code release for paper "Test-Time Training Done Right"
β 500RoboSpatial. Python
β 147ViewSpatial-Bench. [ECCV 2026] ViewSpatial-Bench:Evaluating Multi-perspective Spatial Localization in Vision-Language Models
β 82OmniSpatial. [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
β 89ChatGarment. [CVPR'25] ChatGarment: Garment Estimation, Generation and Editing via Large Language Models
β 168grounded-rl. Python
β 133OpenThinkIMG. OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.
β 399Articulate-Anymesh. C++
β 137UP-VLA. Official PyTorch implementation for ICML 2025 paper: UP-VLA.
β 61HAMSTER_beta. Python
β 62Awesome-3D-Scene-Generation. A curated list of awesome 3D scene generation papers. (arXiv 2505.05474)
β 1.1kViewCrafter. [TPAMI 2025] ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
β 1.6kTesserAct. ICCV 2025 | TesserAct: Learning 4D Embodied World Models
β 404NormalCrafter. [ICCV 2025] NormalCrafter: Learning Temporally Consistent Video Normal from Video Diffusion Priors
β 241hamer. HaMeR: Reconstructing Hands in 3D with Transformers
β 1.1kAether. [ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling
β 6043D-Mem. [CVPR 2025] Source codes for the paper "3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning"
β 270SimplerEnv. Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Google Robot, WidowX+Bridge) (CoRL 2024)
β 1.1kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
β 30kLLaMA-Mesh. Unifying 3D Mesh Generation with Language Models
β 1.2kDimensionX. [ICCV'25]DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
β 1.3kTIP-I2V. [ICCV 2025] TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
β 41RealBasicVSR. Official repository of "Investigating Tradeoffs in Real-World Video Super-Resolution"
β 1.1kEvTexture. [ICML 2024 & TPAMI 2026] EvTexture & EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution
β 1.2kDepthAnyVideo. Depth Any Video with Scalable Synthetic Data (ICLR 2025)
β 518Emu3. Next-Token Prediction is All You Need
β 2.4kmochi. The best OSS video generation models, created by Genmo
β 3.7kAllegro. Allegro is a powerful text-to-video model that generates high-quality videos up to 6 seconds at 15 FPS and 720p resolution from simple text input.
β 1.1kO1-Journey. O1 Replication Journey
β 2kPyramid-Flow. [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
β 3.2kmonst3r. Official Implementation of paper "MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion"
β 1.4kGrounded_3D-LLM. Code&Data for Grounded 3D-LLM with Referent Tokens
β 136PointLLM. [ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
β 1kLLaVA-3D. [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D World
β 388Gear-NeRF. This repository contains the implementation of the paper: "Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling", CVPR 2024 (Highlight)
β 18PuzzleAvatar. [SIGGRAPH Asia 2024] PuzzleAvatar: Assembling 3D Avatars from Personal Albums
β 320SVD_Xtend. Stable Video Diffusion Training Code and Extensions.
β 731aiTour. AI ε¦δΉ δΉζ
β 110super_primitive. [CVPR'24, Demo Track Honourable Mention] SuperPrimitive: Scene Reconstruction at a Primitive Level
β 204CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
β 13kshape-of-motion. Python
β 1.3kawesome-4d-generation. List of papers on 4D Generation.
β 327RAFT. Python
β 4.1kMeshXL. [NeurIPS 2024] MeshXL: Neural Coordinate Field for Generative 3D Foundation Models, a 3D fundamental model for mesh generation
β 339dreamscene4d. [NeurIPS 2024] DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
β 232Open-Sora. Open-Sora: Democratizing Efficient Video Production for All
β 29kMarigold. [CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
β 3.2kgcd. Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis (ECCV 2024 Oral) - Official Implementation
β 290STAG4D. Official Implementation for STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians
β 210EgoThink. [CVPR'24 Highlight] The official code and data for paper "EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models"
β 65HumanVLA. Python
β 135openvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
β 6.7kRapVerse. Code for paper "RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text"
β 18lerobot. π€ LeRobot: Making AI for Robotics more accessible with end-to-end learning
β 26kGPTEval3D. [ CVPR 2024 ] Implementation for "GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation"
β 288llama3. The official Meta Llama 3 GitHub site
β 29kVAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusionπ₯] [scaling laws in visual generationπ] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
β 8.7kJetMoE. Reaching LLaMA2 Performance with 0.1M Dollars
β 986MVDream. Multi-view Diffusion for 3D Generation
β 987PixArt-alpha. PixArt-Ξ±: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
β 3.3ksd-dino. Official Implementation of paper "A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence"
β 357robotic-transformer-pytorch. Implementation of RT1 (Robotic Transformer) in Pytorch
β 454FeatUp. Official code for "FeatUp: A Model-Agnostic Frameworkfor Features at Any Resolution" ICLR 2024
β 1.7khumanoid-bench. Python
β 7793D-VLA. [ICML 2024] 3D-VLA: A 3D Vision-Language-Action Generative World Model
β 630Sandwich. Bidirectional Mapping between Action Physical-Semantic Space
β 34urdfpy. Python parser for URDFs
β 320urdf_to_obj. Repository to extract obj files in world frame from a URDF description
β 21ODICE-Pytorch. official implementation of ODICE
β 19LGM. [ECCV 2024 Oral] LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation.
β 2.1kAwesome-AIGC-3D. A curated list of awesome AIGC 3D papers
β 787Awesome-LLM-3D. Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resources
β 2.2kDepth-Anything. [CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation
β 8.2khiveformer. Python
β 33MultiPLY. Code for MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
β 135octo. Octo is a transformer-based robot policy trained on a diverse mix of 800k robot trajectories.
β 1.7kunified-io-2. Python
β 649MiDaS. Code for robust monocular depth estimation described in "Ranftl et. al., Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer, TPAMI 2022"
β 5.4kRE-OT. Python
β 2LVM. Python
β 1.8kmeshgpt-pytorch. Implementation of MeshGPT, SOTA Mesh generation using Attention, in Pytorch
β 861NExT-GPT. Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
β 3.6kdobb-e. DobbΒ·E: An open-source, general framework for learning household robotic manipulation
β 621BlenderSynth. Synthetic Blender Dataset Production
β 99DeformingThings4D. [ICCV 2021] A dataset of non-rigidly deforming objects.
β 363peft. π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
β 21kRoboGen. A generative and self-guided robotic agent that endlessly propose and master new skills.
β 1.2kblendify. Lightweight Python framework that provides a high-level API for creating and rendering scenes with Blender.
β 865SQA3D. [ICLR 2023] SQA3D for embodied scene understanding and reasoning
β 170PointFlow. PointFlow : 3D Point Cloud Generation with Continuous Normalizing Flows
β 867zero123plus. Code repository for Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model.
β 2.1kGenius-Invokation. δΈε£ε¬ε€εΌΊεε¦δΉ η―ε’
β 39SALMON. Self-Alignment with Principle-Following Reward Models
β 170T2M-GPT. (CVPR 2023) Pytorch implementation of βT2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representationsβ
β 772RLBench. A large-scale benchmark and learning environment.
β 1.8kMiniGPT-5. Official implementation of paper "MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens"
β 868MotionGPT. [NeurIPS 2023] MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs
β 1.9kManiSkill. Manipulation Skill Framework, an open source GPU parallelized robotics simulator and benchmark
β 3.2kcolmap. COLMAP - Structure-from-Motion and Multi-View Stereo
β 12kiQuery. [CVPR 2023] iQuery: Instruments as Queries for Audio-Visual Sound Separation
β 73SportsSloMo. SportsSloMo: A New Benchmark and Baseline Models for Human-centric Video Frame Interpolation, CVPR 2024 (https://arxiv.org/abs/2308.16876)
β 79dreamgaussian. [ICLR 2024 Oral] Generative Gaussian Splatting for Efficient 3D Content Creation
β 4.3kScotty3D. Base code for 15-462/662: Computer Graphics at Carnegie Mellon University
β 554pytorch-openpose. pytorch implementation of openpose including Hand and Body Pose Estimation.
β 2.3kDeformable-3D-Gaussians. [CVPR 2024] Official implementation of "Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction"
β 1.2kgrounded-segment-any-parts. Grounded Segment Anything: From Objects to Parts
β 416DreamLLM. [ICLR 2024 Spotlight] DreamLLM: Synergistic Multimodal Comprehension and Creation
β 462ZoeDepth. Metric depth estimation from a single image
β 2.8ktextured_smplx. Python
β 112MIME. This is the official code for MIME: Human-Aware 3D Scene Generation (CVPR2023)
β 96ACID. ACID: Action-Conditional Implicit Visual Dynamics for Deformable Object Manipulation
β 75objaverse-xl. πͺ Objaverse-XL is a Universe of 10M+ 3D Objects. Contains API Scripts for Downloading and Processing!
β 1.3kmulti3drefer. [ICCV 2023] Multi3DRefer: Grounding Text Description to Multiple 3D Objects
β 98diffhoi_v2. Official Reimplementation of Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction Clips (DiffHOI, ICCV23) https://judyye.github.io/diffhoi-www/
β 36SyncDreamer. [ICLR 2024 Spotlight] SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
β 1kaccelerate. π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
β 9.8kparis. [ICCV 2023] Official implementation of the paper "PARIS: Part-level Reconstruction and Motion Analysis for Articulated Objects"
β 87diff-gaussian-rasterization. Cuda
β 1.5kgaussian-splatting. Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
β 23kSSDNeRF. [ICCV 2023] Single-Stage Diffusion NeRF
β 446DriveLM. [ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering
β 1.3kthreestudio. A unified framework for 3D content generation.
β 7kCoDeF. [CVPR'24 Highlight] Official PyTorch implementation of CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
β 4.8kGPT4Point. [CVPR'24 Highlight] GPT4Point: A Unified Framework for Point-Language Understanding and Generation.
β 443harp. HARP: Personalized Hand Reconstruction from a Monocular RGB Video
β 79Color-NeuS. [3DV 2024] Color-NeuS: Reconstructing Neural Implicit Surfaces with Color
β 146neuralangelo. Official implementation of "Neuralangelo: High-Fidelity Neural Surface Reconstruction" (CVPR 2023)
β 4.6kGrounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
β 18kroformer. Rotary Transformer
β 1.1kLightGlue. LightGlue: Local Feature Matching at Light Speed (ICCV 2023)
β 4.7kdynalang. Code for "Learning to Model the World with Language." ICML 2024 Oral.
β 421