This is your work, valued
vehicle-trajectory-prediction. Behavior Prediction in Autonomous Driving
★ 65Visual-Transformer-Paper-Summary. Summary of Transformer applications for computer vision tasks.
★ 59AdvMix. Official code for our CVPR 2021 paper: "When Human Pose Estimation Meets Robustness: Adversarial Algorithms and Benchmarks".
★ 33Awesome-Monocular-3D.
★ 2Human-Motion-Retargeting-Dataset. A new benchmarking dataset for human motion retargeting
★ 1TripoSR. TripoSR: Fast 3D Object Reconstruction from a Single Image
★ 6.8kKimi-K3. Open Frontier Intelligence
★ 7.5kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93kStep-3.5-Flash. Fast, Sharp & Reliable Agentic Intelligence
★ 2.1ktau-0-vla. This repo is the official implementation of "τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation".
★ 216RPent. RPent: Agentic Infrastructure for the Physical World
★ 188WorldFoundry. Unified World Model Inference & Evaluation Infrastructure
★ 275physical-ai-bench. [CVPR 2026 Oral] PAI-Bench: A Comprehensive Benchmark for Physical AI
★ 91NATTEN. Fast Multi-dimensional Sparse Attention
★ 779SLA. SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention
★ 325ReflectVLN. ReflectVLN: Training Vision-Language Navigation Agents with Self-Reflective Reasoning
★ 7ultralytics. Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
★ 60kCtrl-World. ICLR 2026 Paper: Ctrl-World
★ 538ABot-World. Infinite Interactive World Rollout on a Single Desktop GPU
★ 1.3kflashdreams. high-performance inference and serving library for interactive autoregressive video and world models
★ 431lingbot-video. Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
★ 889NeoVerse. [CVPR 2026 Highlight & Best Paper of VideoWorldModel Workshop] NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
★ 643AgiBotWorldChallengeICRA2026-WorldModelBaseline. Baseline Model Repo for World Model Track in AgiBot World Challenge@ICRA 2026
★ 47T2I-Distill. [Tutorial] Few-Step Distillation for Text-to-Image Generation: A Practical Guide
★ 370RynnWorld-4D. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
★ 74RynnBrain. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Models
★ 846giga-world-1. A Roadmap to Build World Models for Robot Policy Evaluation
★ 844lingbot-vision. Self-supervised learning for spatial perception
★ 885lingbot-vla-v2. From Foundation to Application
★ 682AffordVLA. Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance
★ 16PhysisForcing. PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
★ 113VLAC. VLAC: A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
★ 321giga-world-policy. GigaWorld-Policy: An Efficient Action-Centered World–Action Model
★ 1.4kgiga-world-0. GigaWorld-0: World Models as Data Engine to Empower Embodied AI
★ 1.6kInfinite-Forcing. Infinite-Forcing: Towards Infinite-Long Video Generation
★ 155cosmos-predict2.5. Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the form of video.
★ 1.3kSpargeAttn. [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
★ 1kOne-Forcing. Official codebase for "One-Forcing: Towards Stable One-Step Autoregressive Video Generation"
★ 844dgen. [ICLR 2026] Codebase for paper "Geometry-aware 4D Video Generation for Robot Manipulation"
★ 123unified_video_action. Official PyTorch Implementation of Unified Video Action Model (RSS 2025)
★ 400MotionStream. MotionStream: Real-Time Video Generation with Interactive Motion Controls
★ 575tau-0-wm. Python
★ 300Awesome-Interactive-World-Model. A Comprehensive Survey of Interactive Video World Models
★ 215GLM-V. GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 2.4kGLM-5. GLM-5: From Vibe Coding to Agentic Engineering
★ 6.8kDreamX-World. DreamX-World: A General-Purpose Interactive World Model
★ 738Qwen-VLA. The official repository of Qwen-VLA
★ 725arx5-sdk. C++ and Python controller for ARX5 robot arm developed by @yihuai-gao (@real-stanford)
★ 160ACE-Step-1.5. The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
★ 12kLight-WAM. Codebase for Light-WAM: Efficient World Action Models with State-Fusion Action Decoding.
★ 72audio-intelligence. Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with synthetic captions.
★ 137ABot-Earth-0.5. Generative 3D Earth Model by AMap-cvlab
★ 190robocasa. RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
★ 1.6kcosmos-framework. Our inference and training framework to run on the Cosmos Models
★ 425Open-d4rt. Python
★ 841ReactiveGWM. [Arxiv 2026] ReactiveGWM: Steering NPC in Reactive Game World Models
★ 97UltraVideo. [[NeurIPS 2025] UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
★ 93HumanNet. HumanNet: Scaling Human-centric Video Learning to One Million Hours
★ 278ReVidgen. [ICML 2026🔥]Rethinking Video Generation Model for the Embodied World
★ 89cosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kESI-Bench. Python
★ 119UniVTAC. Jupyter Notebook
★ 121MobileAgent. Mobile-Agent: The Powerful GUI Agent Family
★ 9ksd3.5. Python
★ 1.5kAwesome-Interactive-World-Model. Interactive World Model papers organized by core research challenges.
★ 276Protenix. Toward High-Accuracy Open-Source Biomolecular Structure Prediction.
★ 2kABot-OCR. High-precision document OCR with structured Markdown output
★ 41GE-Sim-V2. Genie Envisioner World Simulator 2.0 (GE-Sim 2.0)
★ 122Warp-as-History. Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
★ 224Dexbotic-RoboChallengeInference. Python
★ 6CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13ktriattention. TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained GPUs.
★ 835Pyramid-Forcing. Official codebase for "Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation"
★ 12Robo3R. Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
★ 69maniparena-repo. Python
★ 55awesome-reliable-robotics. Robotics research demonstrating reliability and robustness in the real world (continuously updated)
★ 160TRELLIS.2. Native and Compact Structured Latents for 3D Generation
★ 9.6kTRELLIS. Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
★ 13kOpenCLAY. CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
★ 980clay. High performance UI layout library in C.
★ 18kMVDream. Multi-view Diffusion for 3D Generation
★ 987DCW. [CVPR 2026] Elucidating the SNR-t Bias of Diffusion Probabilistic Models
★ 120ABot-Claw. Python
★ 200HiDream-O1-Image. Python
★ 1.5kdroid. Distributed Robot Interaction Dataset.
★ 386Open-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12klingbot-va. [RSS 2026] Causal video-action world model for generalist robot control
★ 1.7kDreamDojo. Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
★ 1kForcing-KV. Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".
★ 127minWM. A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
★ 752StereoNav. Jupyter Notebook
★ 29Being-H. Being-H is BeingBeyond's family of human-centric embodied foundation models.
★ 1.1kAR-Diffusion. Python
★ 91vggsfm. VGGSfM: Visual Geometry Grounded Deep Structure From Motion
★ 1.4kHunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kCameraCtrl. Python
★ 658LongCat-Video. Python
★ 5.8kWorldForge. [CVPR 2026 Highlight🎉] Official implementation of WorldForge
★ 169SpaceTimePilot. [CVPR 2026] SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time
★ 123WorldArena. WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
★ 250LDA-1B. [RSS 2026] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
★ 289UniVideo. [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos
★ 544NaVILA. [RSS'25] This repository is the implementation of "NaVILA: Legged Robot Vision-Language-Action Model for Navigation"
★ 676legged-loco. Low-level locomotion policy training in Isaac Lab
★ 447NaVILA-Bench. Vision-Language Navigation Benchmark in Isaac Lab
★ 330cosmos-predict2. Cosmos-Predict2 is a collection of general-purpose world foundation models for Physical AI that can be fine-tuned into customized world models for downstream applications.
★ 794lyra. Project Lyra: Open Generative 3D World Models
★ 2.2kGC-VLN. [CoRL 2025] GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
★ 78WorldMem. [NeurIPS 2025] WorldMem: Long-term Consistent World Simulation with Memory
★ 381Archon. The first open-source harness builder for AI coding. Make AI coding deterministic and repeatable.
★ 23kFrontierNet. [RA-L 2025] FrontierNet: Learning Visual Cues to Explore
★ 169DyWA. [ICCV 2025] DyWA:Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation
★ 93Stable-Video-Infinity. [ICLR 26 Oral] Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
★ 2.5kalpamayo1.5. NVIDIA Alpamayo 1.5 Nano is an open 10B reasoning VLA model for autonomous vehicles with reinforcement-learning enhanced reasoning, navigation guidance, and visual question answering.
★ 345UniWorld. UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
★ 886Nav-R2. Official Implementation of paper: [Nav-R2:Dual‑Relation Reasoning for Generalizable Open‑Vocabulary Object‑Goal Navigation]
★ 20SAPIEN. SAPIEN Embodied AI Platform
★ 801VLA-JEPA. [ECCV 2026] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
★ 511InternVLA-M1. InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
★ 418lerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26krlds. Jupyter Notebook
★ 494robocasa-gr1-tabletop-tasks. Simulation benchmarks of GR1 Tabletop Tasks for GR00T N1
★ 146mujoco. Multi-Joint dynamics with Contact. A general purpose physics simulator.
★ 14kPGSR. [TVCG2024] PGSR: Planar-based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction
★ 1.1kAgentVLN.
★ 179XAgent. An Autonomous LLM Agent for Complex Task Solving
★ 8.5khermes-agent. The agent that grows with you
★ 223kkimi-cli. Kimi Code CLI is your next CLI agent.
★ 11kVEGA-3D. [ECCV 2026] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
★ 421g3D-LF. Official implementation of "g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks" (CVPR'25).
★ 57lerobot-annotate. Lightweight web UI for annotating LeRobot datasets
★ 31ccf-deadlines. ⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
★ 9.2ktab-out. Keep tabs on your tabs. Turn your "New tabs" page into a mission control, so you can close them easily. Built for people who open too many tabs and never close them.
★ 1.7kFast-SAM3D. [ICML 2026] Official Repo for Fast-SAM3D: 3Dfy Anything in Images but Faster
★ 179Qwen3.6. Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.
★ 3.7kHY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kACoT-VLA. [CVPR 2026] Official implementation of "ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models"
★ 224video-prediction-policy. Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations https://video-prediction-policy.github.io
★ 407cosmos-policy. Cosmos Policy
★ 842Reward-Forcing. [CVPR 2026 Highlight] Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
★ 352AnyDepth. AnyDepth: Depth Estimation Made Easy
★ 310q8_kernels. Cuda
★ 82Awesome-Embodied-World-Model. Awesome paper list and repos of the paper "A comprehensive survey of embodied world models".
★ 131InfiniteVGGT. The official implementation of InfiniteVGGT
★ 379GeometryForcing. [ICLR26] Official implementation of Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
★ 214AgiBot-World. [IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
★ 3.1kVid2World. Official repository for "Vid2World: Crafting Video Diffusion Models to Interactive World Models" (ICLR 2026), https://arxiv.org/abs/2505.14357
★ 70dreamzero. Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
★ 2.5kPyramid-Flow. [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
★ 3.2kInfinityStar. [NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
★ 774pillarnext. PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds (CVPR 2023)
★ 247epipolar-dpo. Official repo for: Epipolar Geometry Improves Video Generation Models
★ 96LTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5kLTX-Video. Official repository for LTX-Video
★ 11kMAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.7kNOVA. [ICLR 2025] Autoregressive Video Generation without Vector Quantization
★ 657SkyReels-V3. SkyReels V3: Multimodal Video Generation Model
★ 524cosmos-rl. Cosmos-RL is a flexible and scalable Reinforcement Learning framework specialized for Physical AI applications.
★ 468Sana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6kcausalfusion. Python
★ 197nunchaku. [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
★ 3.9kseraena. WIP Pytorch code for stably training single-step, mode-dropping, deterministic autoencoders
★ 49awesome-diffusion-model-in-rl. A curated list of Diffusion Model in RL resources (continually updated)
★ 1.6kDreamVLA. [NeurIPS 2025] DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
★ 364HPSv2. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
★ 677HPSv3. Official implementation of HPSv3: Towards Wide-Spectrum Human Preference Score (ICCV2025)
★ 331HWM_PLDM. Python
★ 163spirit-v1.5. Spirit-v1.5: A Robotic Foundation Model by Spirit AI
★ 631ABot-PhysWorld. Python
★ 370Infinite-World. [ICML 2026] | Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
★ 195TransArch. Design hardware-friendly model architectures and migrate existing LLMs with minimal performance loss
★ 494Open-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kHunyuan-0.5B. Python
★ 54DiffusionDPO. Code for "Diffusion Model Alignment Using Direct Preference Optimization"
★ 707VideoGPA. [ICML'26] VideoGPA is a self-supervised framework that enhances 3D consistency in Video Diffusion Models.
★ 70VACE. [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
★ 3.9ksam-3d-objects. SAM 3D Objects
★ 7.2kao. PyTorch native quantization and sparsity for training and inference
★ 2.9kcutlass. CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10kAutoFP8. Python
★ 210RaBitQ. The repo has been moved to https://github.com/VectorDB-NTU/RaBitQ-Library. [SIGMOD 2024] RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search
★ 252TurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kTeaCache. Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
★ 1.4kgemm-fp8. High Performance FP8 GEMM Kernels for SM89 and later GPUs.
★ 21SageAttention. [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
★ 3.5kyouclaw. 🦞 Your AI Personal Assistant. An intelligent assistant with memory, skills, and scheduled tasks that truly understands you . Helping you tackle everything in work and life.
★ 722GenRL. Reinforcement Learning Framework for Visual Generation
★ 126Turbo-VAED. [AAAI 2026] Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
★ 134DeepForcing. [ICML 2026] Official implementation of "Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression"
★ 137PackForcing.
★ 255turboquant. TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
★ 1.7kVideoX-Fun. 📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
★ 2.2kAngelSlim. Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
★ 1.5kLightCompress. [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
★ 736ComfyUI-GGUF. GGUF Quantization support for native ComfyUI models
★ 3.9kscannetpp. [ICCV 2023 Oral] ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
★ 396CompACT. Official implementation of "Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model" (CVPR 2026)
★ 64DeepGEMM. DeepGEMM: clean and efficient BLAS kernel library on GPU
★ 7.6kQ-DiT. [CVPR 2025] Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
★ 79RL3DEdit. [ECCV 2026] RL3DEdit
★ 203FastWAM. Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
★ 1.2kWorldStereo. [CVPR 2026] WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories (WorldExpand of HY-World 2.0)
★ 208WildWorld. [ECCV 2026] WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
★ 419URSA. [ICLR 2026] 🐻 Uniform Discrete Diffusion with Metric Path for Video Generation
★ 123Waver. Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.
★ 950