This is your work, valued
BEVerse. The official repository for BEVerse
★ 439OccFormer. [ICCV 2023] OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy Prediction
★ 410MonoFlex. Released code for Objects are Different: Flexible Monocular 3D Object Detection, CVPR21
★ 237GraphAD.
★ 68SimMOD. Implementation of SimMOD: A Simple Baseline for Multi-Camera 3D Object Detection
★ 47Bagel. Open-source unified multimodal model
★ 6.1kgiga-world-policy. GigaWorld-Policy: An Efficient Action-Centered World–Action Model
★ 1.4kEmbodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15kspirit-v1.5. Spirit-v1.5: A Robotic Foundation Model by Spirit AI
★ 631ESAM. [ICLR 2025, Oral] EmbodiedSAM: Online Segment Any 3D Thing in Real Time
★ 634ControlNet. Let us control diffusion models!
★ 34kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kUniVLA. [RSS 2025] Learning to Act Anywhere with Task-centric Latent Actions
★ 1.1kAgiBot-World. [IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
★ 3.1kDepth-Anything-V2. [NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
★ 8.6kopenpi. Python
★ 13kIsaacLab. Unified framework for robot learning built on NVIDIA Isaac Sim
★ 7.8kQwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kcosmos-reason1. Cosmos-Reason1 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
★ 952cosmos-predict1. Cosmos-Predict1 is a collection of general-purpose world foundation models for Physical AI that can be fine-tuned into customized world models for downstream applications.
★ 465RoboticsDiffusionTransformer. RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
★ 1.8kIsaac-GR00T. NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7.7kOpenManus. No fortress, purely open ground. OpenManus is Coming.
★ 58klangchain. The agent engineering platform.
★ 143kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kFlashMLA. FlashMLA: Efficient Multi-head Latent Attention Kernels
★ 13kHE-Drive. HE-Drive: Human-Like End-to-End Driving with Vision Language Models
★ 253OpenEMMA. OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.
★ 946DriveRecon. Python
★ 191GPD. GPD-1: Generative Pre-training for Driving
★ 83Stag. [ICCV 2025] Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model
★ 97tushare. TuShare is a utility for crawling historical data of China stocks
★ 15kIC-Light. More relighting!
★ 8.5kLGM. [ECCV 2024 Oral] LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation.
★ 2.1kOpenCLAY. CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
★ 980SparseDrive. SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
★ 982Vista. [NeurIPS 2024] A Generalizable World Model for Autonomous Driving
★ 889rerun. Visualize, query, and stream to train on multimodal robotics data.
★ 11kneuro-ncap. NeuroNCAP benchmark for end-to-end autonomous driving
★ 253Scale-BEV. Python
★ 54awesome-3D-gaussian-splatting. Curated list of papers and resources focused on 3D Gaussian Splatting, intended to keep pace with the anticipated surge of research in the coming months.
★ 8.8kgenerative-ai-for-beginners. 21 Lessons, Get Started Building with Generative AI
★ 114kDriveAGI. [CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving
★ 801DriveEnvNeRF. [ICRA 2024 Workshop] DriveEnv-NeRF: Exploration of A NeRF-Based Autonomous Driving Environment for Real-World Performance Validation
★ 39GPT4Affectivity. GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
★ 26accelerate. 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
★ 9.8kMGM. Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
★ 3.3kChatSim. [CVPR2024 Highlight] Editable Scene Simulation for Autonomous Driving via LLM-Agent Collaboration
★ 430TEASER-plusplus. A fast and robust point cloud registration library
★ 2.3kPointSAM-for-MixSup. [ICLR 2024] MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object Detection
★ 75pixel-perfect-sfm. Pixel-Perfect Structure-from-Motion with Featuremetric Refinement (ICCV 2021, Best Student Paper Award)
★ 1.5kForge_VFM4AD. A comprehensive survey of forging vision foundation models for autonomous driving, including challenges, methodologies, and opportunities.
★ 272Generalizable-BEV. Python
★ 149RoMe. Python
★ 279textgen. Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.
★ 48kOccWorld. [ECCV 2024] 3D World Model for Autonomous Driving
★ 574Lidar_AI_Solution. A project demonstrating Lidar related AI solutions, including three GPU accelerated Lidar/camera DL networks (PointPillars, CenterPoint, BEVFusion) and the related libs (cuPCL, 3D SparseConvolution, YUV2RGB, cuOSD,).
★ 1.9kSelfOcc. [CVPR 2024] SelfOcc: Self-Supervised Vision-Based 3D Occupancy Prediction
★ 387EmerNeRF. PyTorch Implementation of EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision
★ 639imap. High-resolution map (OpenDrive\Apollo) visualization and conversion tools
★ 262Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kwaymax. A JAX-based simulator for autonomous driving research.
★ 1.1kplanTF. [ICRA'2024] Rethinking Imitation-based Planner for Autonomous Driving
★ 372ADAPT. This repository is an official implementation of ADAPT: Action-aware Driving Caption Transformer, accepted by ICRA 2023.
★ 421llama. Inference code for Llama models
★ 60kCausalHTP. The official PyTorch code implementation of "Human Trajectory Prediction via Counterfactual Analysis" in ICCV 2021.
★ 78DriveDreamer. [ECCV 2024] DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
★ 587scenarionet. ScenarioNet: Scalable Traffic Scenario Management System for Autonomous Driving
★ 299InsMapper. [ECCV 2024] InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping
★ 48VMA. A general map auto annotation framework based on MapTR, with high flexibility in terms of spatial scale and element type
★ 325TAP. [ICCV 2023] Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud Models
★ 43neuralsim. neuralsim: 3D surface reconstruction and simulation based on 3D neural rendering.
★ 671AD-MLP. Python
★ 239End-to-end-Autonomous-Driving. [IEEE T-PAMI 2024] All you need for End-to-end Autonomous Driving
★ 3.7konnx-modifier. A tool to modify ONNX models in a visualization fashion, based on Netron and Flask.
★ 1.6kDFML. Official implementation of Deep Factorized Metric Learning.
★ 20VisualGLM-6B. Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型
★ 4.2kMake-A-Protagonist. Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
★ 322SurroundOcc. [ICCV 2023] SurroundOcc: Multi-camera 3D Occupancy Prediction for Autonomous Driving
★ 1.1kOccFormer. [ICCV 2023] OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy Prediction
★ 411segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kflex-dm. [CVPR 2023 highlight] Towards Flexible Multi-modal Document Models
★ 594d-occ-forecasting. CVPR 2023: Official code for `Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting'
★ 250OpenOccupancy. [ICCV 2023] OpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy Perception
★ 741OccDepth. Maybe the first academic open work on stereo 3D SSC method with vision-only input.
★ 319occupancy-for-nuscenes. 3D occupancy
★ 409CVPR2023-3D-Occupancy-Prediction. CVPR2023-Occupancy-Prediction-Challenge
★ 878ASAP. [CVPR 2023] Are We Ready for Vision-Centric Driving Streaming Perception? The ASAP Benchmark
★ 88TPVFormer. [CVPR 2023] An academic alternative to Tesla's occupancy network for autonomous driving.
★ 1.4kOPERA. Official implementation of "OPERA: Omni-Supervised Representation Learning with Hierarchical Supervisions"
★ 34OpenCC. Automatic driving long tail / corner cases scenarios dataset (Anomaly detection)
★ 114pylot. Modular autonomous driving platform running on the CARLA simulator and real-world vehicles.
★ 535SOLOFusion. Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection
★ 271synthehicle. [WACVW 2023] A massive synthetic dataset for 3D multi-target multi-camera tracking and segmentation.
★ 54Learning-Deep-Learning. Paper reading notes on Deep Learning and Machine Learning
★ 1.3kMapTR. [ICLR'23 Spotlight & ECCV'24 & IJCV'24] MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
★ 1.5kIDML. [TPAMI 2023] Official implementation of Introspective Deep Metric Learning.
★ 55standard-readme. A standard style for README files
★ 6.3kDeepInteraction. [NeurIPS 2022 & TPAMI 2025] DeepInteraction: 3D Object Detection via Modality Interaction
★ 267MulimgViewer. MulimgViewer is a multi-image viewer that can open multiple images in one interface, which is convenient for image comparison and image stitching.
★ 1.4kDEVIANT. [ECCV 2022] Official PyTorch Code of DEVIANT: Depth Equivariant Network for Monocular 3D Object Detection
★ 226omni3d. Code release for "Omni3D A Large Benchmark and Model for 3D Object Detection in the Wild"
★ 854BEVerse. The official repository for BEVerse
★ 439SurroundDepth. [CoRL 2022] SurroundDepth: Entangling Surrounding Views for Self-Supervised Multi-Camera Depth Estimation
★ 293HDMapNet. Python
★ 764BEVDet. Code base of the BEVDet series .
★ 1.8kMonoScene. [CVPR 2022] "MonoScene: Monocular 3D Semantic Scene Completion": 3D Semantic Occupancy Prediction from a single image
★ 818SAFNet. [IROS 2021] Implementation of "Similarity-Aware Fusion Network for 3D Semantic Segmentation"
★ 23object-detection-usages. The brief implementation and using examples of object detection usages like, IoU, NMS, soft-NMS, SmoothL1、IoU loss、GIoU loss、 DIoU loss、CIoU loss, cross-entropy、focal-loss、GHM, AP/MAP and so on by Pytorch.
★ 200MonoFlex. Released code for Objects are Different: Flexible Monocular 3D Object Detection, CVPR21
★ 237leetcode. LeetCode Solutions: A Record of My Problem Solving Journey.( leetcode题解,记录自己的leetcode解题之路。)
★ 56kCaDDN. Categorical Depth Distribution Network for Monocular 3D Object Detection (CVPR 2021 Oral)
★ 402