This is your work, valued
PhDing in Vision and Stuff
salt. Segment Anything Labelling Tool
★ 1kroadwork-dataset. Repository that contains simple scripts to use ROADWork dataset.
★ 56boost-document. GSoC'15 Proposal for new Boost.Document Library
★ 4anuragxel.github.io. Web Page, eh ?
★ 3wiki-doc-classification. Python
★ 2kumba-server. Python
★ 1data-oriented-cluster-hashing. This project is an implementation of Cluster-Based Data Oriented Hashing (Chafik et al, DSAA. 2015)
★ 1dex-lang. Research language for array processing in the Haskell/ML family
★ 1.7kspark. :sparkles: An advanced 3D Gaussian Splatting renderer for THREE.js
★ 3.5kphyco. Checkpoint and evaluation code for PhyCo [CVPR 2026]
★ 8gluemap. GLUEMAP: Global Structure-from-Motion Meets Feedforward Reconstruction
★ 323OneVL_training. Python
★ 25prix. PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving
★ 42FrameCrafter. [ECCV 2026] FrameCrafter: Novel View Synthesis as Video Completion
★ 79street_crafter. [CVPR 2025] StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
★ 329OpenVO. Implementation of Open-World Visual Odometry with Temporal Dynamics Awareness (CVPR'26)
★ 114onevl. Python
★ 456starVLA. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
★ 3.3kvggt-omega. [CVPR 2026 Oral] VGGT Omega
★ 3.8kdspy. DSPy: The framework for programming—not prompting—language models
★ 36kpy123d. 123D: A Unified Library for Multi-Modal Autonomous Driving Data
★ 375ClickHouse. ClickHouse® is a real-time analytics database management system
★ 49kcolmap. COLMAP - Structure-from-Motion and Multi-View Stereo
★ 12kWorkZone3D. A repository for the multimodal WorkZone3D dataset
★ 5tuna-2. Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
★ 739fail2drive. Fail2Drive: Benchmarking Closed-Loop Driving Generalization
★ 157WildDet3D. Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild
★ 603waymax. A JAX-based simulator for autonomous driving research.
★ 1.1kLoMa. [ECCV 2026 Oral] LoMa: Local Feature Matching Revisited
★ 367much-ado-about-noising. A optimized PyTorch framework for behavior cloning with flow related generative models.
★ 287NuPlanQA.
★ 20reachy_mini. Reachy Mini's SDK
★ 1.4kMOTIP. [CVPR 2025] Multiple Object Tracking as ID Prediction
★ 548OLMo-core. PyTorch building blocks for the OLMo ecosystem
★ 1.4kLTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5ksam-audio. The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 3.6kcosmos-transfer2.5. Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.
★ 711Cosmos-Drive-Dreams. Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
★ 518pyvrs. Python interface for https//github.com/facebookresearch/vrs.
★ 58attention-map-diffusers. 🚀 Cross attention map tools for huggingface/diffusers
★ 411ConceptAttention. ConceptAttention: A method for interpreting multi-modal diffusion transformers.
★ 461sam3. The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 11kOpenEMMA. OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.
★ 946CaRL. [CoRL 2025] CaRL: Learning Scalable Planning Policies with Simple Rewards
★ 157InteractiveClosedLoop.
★ 20alpasim. AlpaSim is an open-source autonomous vehicle simulation platform designed for development and testing of end-to-end AV policies
★ 1.1kmpsfm. MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion (CVPR 2025)
★ 525interPlan. Python
★ 131ryugraph. Ryu, a fork of Kuzu, is an Embedded Property Graph Database built for speed with vector search and full-text search built in. Implements Cypher.
★ 139torchax. torchax is a PyTorch frontend for JAX. It gives JAX the ability to author JAX programs using familiar PyTorch syntax. It also provides JAX-Pytorch interoperability, meaning, one can mix JAX & Pytorch syntax together when authoring ML programs, and run it in every hardware JAX can run.
★ 233GEN3C. [CVPR 2025 Highlight] GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
★ 1.4kopenpi. Python
★ 13kDepth-Anything-V2. [NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
★ 8.6kMoGe. [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
★ 2.7klqrax. GPU-friendly, auto-differentiable LQR solver with JAX.
★ 222xinfer. Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.
★ 295gpt-oss. gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
★ 20kharmony. Renderer for the harmony response format to be used with gpt-oss
★ 4.5kplanTF. [ICRA'2024] Rethinking Imitation-based Planner for Autonomous Driving
★ 373pluto. PLUTO: Push the Limit of Imitation Learning-based Planning for Autonomous Driving
★ 609safe-sim. Python
★ 63nano-vllm. Nano vLLM
★ 15kDiffusion-Planner. [ICLR 2025 Oral] The official implementation of "Diffusion-Based Planning for Autonomous Driving with Flexible Guidance"
★ 1kpyslam. pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras. It provides a broad set of modern local and global feature extractors, multiple loop-closure strategies, a volumetric reconstruction module, integrated depth-prediction models, and semantic segmentation capabilities for enhanced scene understanding.
★ 3.4kPrior-Depth-Anything. Python
★ 520fastmap. A fast and simple structure from motion pipeline written in Pytorch.
★ 282Patch_Scaling. Official implementation of Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
★ 25aerial-megadepth. [CVPR 2025] AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
★ 189infinigen. Infinite Photorealistic Worlds using Procedural Generation
★ 7.2kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kimm. Official implementation of Inductive Moment Matching
★ 585MegaLoc. An image retrieval model for any localization task
★ 265gpudrive. 1 million FPS multi-agent driving simulator
★ 609samurai. Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
★ 7.1kmamba. The Fast Cross-Platform Package Manager
★ 8.1kUniDepth. Universal Monocular Metric Depth Estimation
★ 1.2kInstance-Warp. [WACV 2025] Instance-Level Image Warping for Domain Adaptation
★ 4inverse_painting. Inverse Painting: Reconstructing The Painting Process (SIGGRAPH ASIA 2024)
★ 198rope-vit. [ECCV 2024] Official PyTorch implementation of RoPE-ViT "Rotary Position Embedding for Vision Transformer"
★ 467detrex. detrex is a research platform for DETR-based object detection, segmentation, pose estimation and other visual recognition tasks.
★ 2.3kvismatch. Wrapper of 50+ image matching models with a unified interface
★ 899ml-sigmoid-attention. Python
★ 310droid_metric. run DROID-SLAM with Metric3D to improve monocular performance
★ 185drivestudio. A 3DGS framework for omni urban scene reconstruction and simulation.
★ 1.2kmesh-data-synthesizer. Uses Unreal Engine & Cesium to generate large synthetic dataset from 3D meshes. Enables machine learning tasks like Visual Place Recognition read more in our paper on this: https://meshvpr.github.io
★ 37SCLIP. Official implementation of SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
★ 193glasbey. Algorithmically create or extend categorical colour palettes
★ 237rerun. Visualize, query, and stream to train on multimodal robotics data.
★ 11kGrendel-GS. [ICLR 2025 Oral] On Scaling Up 3D Gaussian Splatting Training
★ 680Stable-DINO. [ICCV 2023] Official implementation of the paper "Detection Transformer with Stable Matching"
★ 243glomap. [DEPRECATED] GLOMAP - Global Structured-from-Motion Revisited
★ 2.4keinx. Universal Notation for Tensor Operations in Python
★ 520assetto_corsa_gym. Assetto Corsa OpenAI Gym Environment
★ 199x360dataset-kit. Python
★ 38Marigold. [CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
★ 3.2ksigma-gpt. σ-GPT: A New Approach to Autoregressive Models
★ 77mixed-resolution-vit. Python
★ 58roadwork-dataset. Repository that contains simple scripts to use ROADWork dataset.
★ 56RPG-DiffusionMaster. [ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
★ 1.8kContext-Cluster. [ICLR 2023 Oral] Image as Set of Points
★ 575scaling_on_scales. When do we not need larger vision models?
★ 419pylot. Modular autonomous driving platform running on the CARLA simulator and real-world vehicles.
★ 535cad. Content-Adaptive Downsampling in Convolutional Neural Networks (CVPR 2023 Workshop on Efficient Deep Learning for Computer Vision)
★ 31gaussian_surfels. [SIGGRAPH'24] Implementations for "High-quality Surface Reconstruction using Gaussian Surfels".
★ 682DROID-SLAM. Python
★ 2.6kMS-DOS. The original sources of MS-DOS 1.25, 2.0, and 4.0 for reference purposes
★ 32kassignment4. 3D Gaussian Splatting and Diffusion Guided Optimization
★ 38smalldiffusion. Simple and readable code for training and sampling from diffusion models
★ 777SKS-Homography. Implementation of PAMI 2025 paper "Fast and Interpretable 2D Homography Decomposition: Similarity-Kernel-Similarity and Affine-Core-Affine Transformations"
★ 19OpenTrafficCam3D. Source code for paper Toward Planet-Wide Traffic Camera Calibration (WACV 2024)
★ 17dust3r. DUSt3R: Geometric 3D Vision Made Easy
★ 7.3kvllm-code-harness. Run code inference-only benchmarks quickly using vLLM
★ 9assignment3. Volume Rendering, Neural Radiance Fields, Neural Surfaces
★ 25doppelgangers. Doppelgangers: Learning to Disambiguate Images of Similar Structures
★ 206GPT-Driver. Learning to Drive with GPT
★ 303StreamingForecasting. [IROS 2023] "Streaming Motion Forecasting for Autonomous Driving"
★ 412pcnet. Python
★ 97OpenSeeD. [ICCV 2023] Official implementation of the paper "A Simple Framework for Open-Vocabulary Segmentation and Detection"
★ 762img2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4.4kChanakya. Learning Runtime Decisions for Adaptive Real-Time Perception
★ 9voicecamo. Code for the paper Real-Time Neural Voice Camouflage
★ 28gaussian-splatting. Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
★ 23ktwo-plane-prior. Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object Detection
★ 6DeepImageBlending. This is a Pytorch implementation of deep image blending
★ 471hiera. Hiera: A fast, powerful, and simple hierarchical vision transformer.
★ 1.1klzu. Code for Learning to Zoom and Unzoom (CVPR 2023)
★ 47submitit. Python 3.8+ toolbox for submitting jobs to Slurm
★ 1.6kSAM-Tool. 利用Segment Anything(SAM)模型进行快速标注
★ 237salt. Segment Anything Labelling Tool
★ 1kpyboreas. Devkit for the Boreas autonomous driving dataset.
★ 121HAMS. Code for Automated License Testing from the HAMS group @ Microsoft Research, India
★ 8TaskMatrix. Python
★ 34kllama.cpp. LLM inference in C/C++
★ 122kslahmr. Python
★ 534yolov7. Implementation of paper - YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
★ 14kllama. Inference code for Llama models
★ 60karxiv-submission-sanitizer-flattener. Simple Python scripts to clean up and flatten ArXiv LaTeX submissions.
★ 69theseus. A library for differentiable nonlinear optimization
★ 2.1kimagenetx. understanding model mistakes with human annotations
★ 105kornia. 🐍 Geometric Computer Vision Library for Spatial AI
★ 11kVanishingPoint_HoughTransform_GaussianSphere. Official implementation: Deep vanishing point detection: Geometric priors make dataset variations vanish, CVPR'22.
★ 115av2-api. Argoverse 2: Next generation datasets for self-driving perception and forecasting.
★ 413pytorch3d. PyTorch3D is FAIR's library of reusable components for deep learning with 3D data
★ 9.9kDetic. Code release for "Detecting Twenty-thousand Classes using Image-level Supervision".
★ 2kmmcv. OpenMMLab Computer Vision Foundation
★ 6.5kmmtracking. OpenMMLab Video Perception Toolbox. It supports Video Object Detection (VID), Multiple Object Tracking (MOT), Single Object Tracking (SOT), Video Instance Segmentation (VIS) with a unified framework.
★ 3.9kon-demand-fork. On-demand-fork
★ 33cvxpylayers. Differentiable convex optimization layers
★ 2.1kmmdetection. OpenMMLab Detection Toolbox and Benchmark
★ 33kVerified-Telemetry. Azure Verified Telemetry for IoT is a state-of-the-art solution to seamlessly determine the health of the sensor in real-time.
★ 20HTML4Vision. A simple HTML visualization tool for computer vision research :hammer_and_wrench:
★ 251ApproxDet. Python
★ 12UniverseNet. USB: Universal-Scale Object Detection Benchmark (BMVC 2022)
★ 431tide. A General Toolbox for Identifying Object Detection Errors
★ 741cocoapi. COCO API - Dataset @ http://cocodataset.org/
★ 46jetson_stats. 📊 Simple package for monitoring and control your NVIDIA Jetson [Orin, Xavier, Nano, TX] series
★ 2.6ksAP. Code for Towards Streaming Perception (ECCV 2020) :car:
★ 101jsondaora. Interoperates dataclasses and TypedDict annotations with json objects for python
★ 39jsonschema-typed. Use JSON Schema for type checking in Python
★ 41serve. Serve, optimize and scale PyTorch models in production
★ 4.4kandroid_rinex. This repository contains a python script that converts logs from Google's GNSS measurement tools to RINEX
★ 104VoTT. Visual Object Tagging Tool: An electron app for building end to end Object Detection Models from Images and Videos.
★ 4.4kawesome-computer-vision-in-sports. A comprehensive list of papers on computer vision in sports
★ 103pytorch_divcolor. Diverse Colorization in Torch
★ 19NeuralNetwork-Viterbi. Python
★ 54tcfpn-isba. Weakly-Supervised Action Segmentation with Iterative Soft Boundary Assignment (CVPR 2018)
★ 41labelImg. LabelImg is now part of the Label Studio community. The popular image annotation tool created by Tzutalin is no longer actively being developed, but you can check out Label Studio, the open source data labeling tool for images, text, hypertext, audio, video and time-series data.
★ 25kkeras-js. Run Keras models in the browser, with GPU support using WebGL
★ 5ktrytond. Mirror of trytond
★ 166netvlad. NetVLAD: CNN architecture for weakly supervised place recognition
★ 608image-forensics. Java
★ 239cvpr2015. Jupyter Notebook
★ 869imagenet-multiGPU.torch. an imagenet example in torch.
★ 407Hindi-Latex-Template. TeX
★ 3visualindex. A simple demo of visual object matching using VLFeat
★ 35CNTK. Microsoft Cognitive Toolkit (CNTK), an open source deep-learning toolkit
★ 18kEigenLibSVM. A wrapper for LibSVM that lets you train SVM's directly on Eigen library matrices in C++
★ 109CanvasInput. HTML5 Canvas Text Input
★ 519PureDarwin. Darwin is the Open Source core of macOS, and PureDarwin is a community project to extend Darwin into a complete, usable operating system.
★ 2.6kAutomated-Essay-Grading. Python
★ 2electron. :electron: Build cross-platform desktop apps with JavaScript, HTML, and CSS
★ 122krapt. Robots Are People Too, a platformer in HTML5
★ 224console2-solarized. The Solarized theme ported to Console2
★ 120Cello. Higher level programming in C
★ 7.1ktodo. My Terminal todo app
★ 4gdc. GNU D Compiler
★ 360