This is your work, valued
Computer Vision, Machine Learning and Anomaly Detection
Object-Detection-Metrics. Most popular metrics used to evaluate object detection algorithms.
★ 5.1kreview_object_detection_metrics. Object Detection Metrics. 14 object detection metrics: mean Average Precision (mAP), Average Recall (AR), Spatio-Temporal Tube Average Precision (STT-AP). This project supports different bounding box formats as in COCO, PASCAL, Imagenet, etc.
★ 1.2kdarknet. Useful functionalities added on the original darknet public repository.
★ 38Deep-Learning-Topics. Jupyter Notebook
★ 20DeepLearning-VDAO. Here are some of the results of my experiments applying Deep Learning for object detection.
★ 19TCF-LMO. TCF-LMO is a network made with dedicated modules to process videos and identify the presence of anomalies in frames. It is composed by: dissimilarity model; a differentiable morphology module; temporal consistency; and classification module.
★ 12image_dedupe. Python
★ 8tennis_autodistill. Jupyter Notebook
★ 6telepic. A lightweight web application for remotely viewing images from a remote computer through a web browser. 🖼️
★ 6awesome-deep-learning. A curated list of awesome Deep Learning tutorials, projects and communities.
★ 5Pos-Palmas-Modulo-Xamarin. Aulas do módulo de Xamarin do curso de Pós-Graduação em Desenvolvimento de Software para Dispositivos Móveis (Católica Palmas-TO)
★ 3Pos-Palmas-Modulo-CSharp. Aulas do módulo de C# do curso de Pós-Graduação em Desenvolvimento de Software para Dispositivos Móveis (Católica Palmas-TO)
★ 3spatial-temporal-action-detection. Python
★ 2vdao-anomaly. Code for "Moving-camera Video Surveillance in Cluttered Environments using Deep Features" (ICIP 2018)
★ 23W. Timely detections for more proactive and effective actions in offshore oil wells!
★ 2supervision. We write your reusable computer vision tools. 💜
★ 1transformers. 🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
★ 1create_object_detector_leaderboard. Repo used to generate the results for HuggingFace's Object Detector Leaderboard
★ 1darknet-1. Convolutional Neural Networks
★ 1RoboticsAcademy. Learn Robotics with JdeRobot
★ 1turbovec. A vector index built on TurboQuant, written in Rust with Python bindings
★ 14kPixelRAG. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
★ 7.9kVision-Agents. Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
★ 8kpyts. A Python package for time series classification
★ 1.9kOmniShotCut. OmniShotCut is a sensitive and more informative SoTA on Shot Boundary Detection task.
★ 263skills. Give your agents the power of the Hugging Face ecosystem
★ 11kINSID3. [CVPR 2026 Oral] "INSID3: Training-Free In-Context Segmentation with DINOv3"
★ 700EUPE. Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized experts across multiple task domains.
★ 690claw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195khf-mount. Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.
★ 772gim. GIM: Learning Generalizable Image Matcher From Internet Videos (ICLR 2024 Spotlight)
★ 900vismatch. Wrapper of 50+ image matching models with a unified interface
★ 900skills. Public repository for Agent Skills
★ 165kijepa. Official codebase for I-JEPA, the Image-based Joint-Embedding Predictive Architecture. First outlined in the CVPR paper, "Self-supervised learning from images with a joint-embedding predictive architecture."
★ 3.5kjepa. PyTorch code and models for V-JEPA self-supervised learning from video.
★ 4.1kAction100M. A Large-scale Video Action Dataset
★ 483uniface. UniFace: A Unified Face Analysis Library for Python | Detection, alignment, landmarks, recognition, parsing, gaze, attributes and anti-spoofing under one API.
★ 779vllm-omni. A framework for efficient model inference with omni-modality models
★ 5.7kbuild-your-own-x. Master programming by recreating your favorite technologies from scratch.
★ 533kSANSA. [NeurIPS 2025 Spotlight] "SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation."
★ 203sam3. The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 11kFeatUp. Official code for "FeatUp: A Model-Agnostic Frameworkfor Features at Any Resolution" ICLR 2024
★ 1.7klivestream_saver. Monitor a Youtube channel and download live-streams from the first segment
★ 397Rex-Omni. [CVPR2026] Detect Anything via Next Point Prediction
★ 1.5ktldw. Too Long, Didn't Watch: End-to-End Rolling Summarizer of Long Videos
★ 375streaming-vlm. StreamingVLM: Real-Time Understanding for Infinite Video Streams
★ 1ksegdino. SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
★ 288mcp. Browser MCP is a Model Context Provider (MCP) server that allows AI applications to control your browser
★ 6.9kchrome-devtools-mcp. Chrome DevTools for coding agents
★ 48kDINOv3_Distillation_YOLO-pose. This project provides a complete pipeline to pre-train the backbone of a custom YOLOv11 pose estimation model using knowledge distillation from a powerful DINOv3 vision foundation model.
★ 58anamorpher. image scaling attacks for multi-modal prompt injection
★ 1.1kdinov3. Reference PyTorch implementation and models for DINOv3
★ 11ktracklab. A Modular End-to-End Tracking Framework for Research and Development 🎯🔬
★ 243truss-examples. Examples of models deployable with Truss
★ 228truss. The simplest way to serve AI/ML models in production
★ 1.2kEdgeTAM. [CVPR 2025] Official PyTorch implementation of "EdgeTAM: On-Device Track Anything Model"
★ 952vipe. ViPE: Video Pose Engine for Geometric 3D Perception
★ 2.1ktennis-tracking. Open-source Monocular Python HawkEye for Tennis
★ 699asm-lessons. FFmpeg Assembly Language Lessons
★ 12koptimum-onnx. 🤗 Optimum ONNX: Export your model to ONNX and run inference with ONNX Runtime
★ 158mmaction2. OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark
★ 5.1krf-detr. RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
★ 8.8kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kdeepstream_python_apps. DeepStream SDK Python bindings and sample applications
★ 1.9kvidgear. A High-performance cross-platform Video Processing Python framework powerpacked with unique trailblazing features :fire:
★ 3.7kdatumaro. Dataset Management Framework, a Python library and a CLI tool to build, analyze and manage Computer Vision datasets.
★ 55datumaro. Dataset Management Framework, a Python library and a CLI tool to build, analyze and manage Computer Vision datasets.
★ 683dl-4-tsc. Deep Learning for Time Series Classification
★ 1.7kresponses.js. A lightweight express.js server implementing OpenAI’s Responses API, built on top of Chat Completions, powered by Hugging Face Inference Providers.
★ 234accelerate. 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
★ 9.8kMarigold. [CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
★ 2mcp-images. ## MCP-Images Looking for a powerful image processing server? MCP Server-Image provides enterprise-grade image handling with just a few lines of code. Perfect for AI applications, web services, and data processing pipelines. [Get Started](#installation) | [Support Us](https://www.buymeacoffee.com/blazzmocompany)
★ 20pytorch-grad-cam. Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
★ 13kgemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106knotebooks. 250+ Fine-tuning & RL Notebooks for text, vision, audio, embedding, TTS models.
★ 5.5kclaude-desktop. Claude Desktop for Debian-based Linux distributions
★ 129jan. Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
★ 44kMcByte. [CVPRW 2025] McByte - tracking in sports without training (No Train Yet Gain)
★ 108mcp-code-review-server. A MCP server for code reviews
★ 34ShowUI. [CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.
★ 1.9kagenticSeek. Fully Local Manus AI. No APIs, No $200 monthly bills. Enjoy an autonomous agent that thinks, browses the web, and code for the sole cost of electricity.
★ 27ksmol-vision. Recipes for shrinking, optimizing, customizing cutting edge vision models. 💜
★ 2kmplbasketball. Basketball plotting library for use with matplotlib
★ 91timesfm. TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
★ 27ktriton. Development repository for the Triton language and compiler
★ 20kkernels. Build compute kernels and load them from the Hub.
★ 718RADIO. Official repository for "AM-RADIO: Reduce All Domains Into One"
★ 1.9kact. Run your GitHub Actions locally 🚀
★ 71ktrackers. Trackers gives you clean, modular re-implementations of leading multi-object tracking algorithms released under the permissive Apache 2.0 license. You combine them with any detection model you already use.
★ 3.6kmcp-use. The fullstack MCP framework to develop MCP Apps for ChatGPT / Claude & MCP Servers for AI Agents.
★ 10kuniversal-intelligence. ◉ Universal Intelligence: AI made simple.
★ 61nerfstudio. A collaboration friendly studio for NeRFs
★ 12kawesome-mcp-servers. A collection of MCP servers.
★ 92kdeephar. Deep human action recognition and pose estimation
★ 423LW-DETR. This repository is an official implementation of the paper "LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection".
★ 505create-python-server. Create a Python MCP server
★ 476claude-desktop-debian. Claude Desktop for Linux
★ 5.3kmcp-hfspace. MCP Server to Use HuggingFace spaces, easy configuration and Claude Desktop mode.
★ 388telepic. A lightweight web application for remotely viewing images from a remote computer through a web browser. 🖼️
★ 6image_dedupe. Python
★ 8DarkIR. CVPR 2025 DarkIR: Robust Low-Light Image Restoration - State of the art low light deblurring. NTIRE 2025 Best Method. [Official PyTorch Implementation]
★ 352TennisCourtDetector. Deep learning network for detecting tennis court
★ 269servers. Model Context Protocol Servers
★ 89kawesome-mcp-servers. Awesome MCP Servers - A curated list of Model Context Protocol servers
★ 5.7kyoloe. YOLOE: Real-Time Seeing Anything [ICCV 2025]
★ 2.2kRexSeek. [ICCV2025] Referring any person or objects given a natural language description. Code base for RexSeek and HumanRef Benchmark
★ 184FindTrack. [ICCVW 2025] Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation
★ 82RGVI. [AAAI 2025] Elevating Flow-Guided Video Inpainting with Reference Generation
★ 96OpenManus. No fortress, purely open ground. OpenManus is Coming.
★ 58kpython-sdk. The official Python SDK for Model Context Protocol servers and clients
★ 24kDistill-Any-Depth. The repo for "Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator"
★ 693gpu-fryer. Where GPUs get cooked 👩🍳🔥
★ 401nerfies.github.io. JavaScript
★ 4.3kVILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8kmaestro. streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
★ 2.7kyolov12. [NeurIPS 2025] YOLOv12: Attention-Centric Real-Time Object Detectors
★ 2.9kpaperpulse. Daily summaries of Arxiv papers
★ 7tennis_autodistill. Jupyter Notebook
★ 6Object-Detection-and-Tracking. Multi-Object Tracking via DeepSORT
★ 2kopen-r1. Fully open reproduction of DeepSeek-R1
★ 26kYOLO. An MIT License of YOLOv9, YOLOv7, YOLO-RD
★ 1.7kGroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
★ 10kDynOMo. Official code of DynOMo: Online Point Tracking by Dynamic Online Monocular Gaussian Reconstruction (3DV 2025))
★ 174browser-use. 🌐 Make websites accessible for AI agents. Automate tasks online with ease.
★ 107kDeepSeek-R1.
★ 92kautodistill. Images to inference with no labeling (use foundation models to train supervised models).
★ 2.8ksafetensors. Simple, safe way to store and distribute tensors
★ 3.8ktorchMoji. 😇A pyTorch implementation of the DeepMoji model: state-of-the-art deep learning model for analyzing sentiment, emotion, sarcasm etc
★ 922segmentation_models.pytorch. Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.
★ 12kvlms-zero-to-hero. This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
★ 1.2ksglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kDINO-X-API. DINO-X: The World's Top-Performing Vision Model for Open-World Object Detection and Understanding
★ 1.4ktrl. Train transformer language models with reinforcement learning.
★ 19kamazon-sagemaker-multiple-object-tracking. Python
★ 15Liger-Kernel. Efficient Triton Kernels for LLM Training
★ 6.5kbtop. A monitor of resources
★ 34kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kpillow-jpegxl-plugin. Pillow plugin for JPEG-XL, using Rust for bindings.
★ 61tvcalib. Python
★ 49smol-course. A course on aligning smol models.
★ 6.7kimg2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4.4kopen_clip. An open source implementation of CLIP.
★ 14ksportsfield_release. Code release for WACV 2020, "Optimizing Through Learned Errors for Accurate Sports Field Registration"
★ 77RollingDepth. [CVPR 2025] RollingDepth: Video Depth without Video Models
★ 610EfficientTAM. Efficient Track Anything
★ 815boxmot. BoxMOT: Pluggable Python and C++ SOTA multi-object tracking modules with support for axis-aligned and oriented bounding boxes
★ 8.3kRectlabel-support. RectLabel is an offline image annotation tool for object detection and segmentation.
★ 552meshgen. Use AI Agents directly in Blender.
★ 900ssd.pytorch. A PyTorch Implementation of Single Shot MultiBox Detector
★ 5.2kcaffe. Caffe: a fast open framework for deep learning.
★ 4.8ksamurai. Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
★ 7.1kLitServe. A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.
★ 3.9kbeartype. Unbearably fast near-real-time pure-Python runtime-static type-checker.
★ 3.5kUnSAM. [NeurIPS 2024] Code release for "Segment Anything without Supervision"
★ 503pr-agent. 🚀 PR Agent: The Original Open-Source PR Reviewer. This project is not the Qodo free tier.
★ 12kPitchGeometry. Detect keypoints at a football pitch
★ 26homographynet. Implementation of Deep Image Homography Estimation (HomographyNet) by DeTone, Malisiewicz, and Rabinovich
★ 47smart-commit. Smart commit messages
★ 18jupyter-bbox-widget. A Jupyter widget for annotating images with bounding boxes
★ 143segment-geospatial. A Python package for segmenting geospatial data with the Segment Anything Model (SAM)
★ 4.1kChangeMamba. [IEEE TGRS 2024] ChangeMamba: Remote Sensing Change Detection Based on Spatio-Temporal State Space Model
★ 634awesome-remote-sensing-change-detection. A comprehensive and up-to-date compilation of datasets, tools, methods, review papers, and competitions for remote sensing change detection.
★ 2.3kSynthMoCap. SynthMoCap Datasets
★ 472Jobs_Applier_AI_Agent_AIHawk. AIHawk aims to easy job hunt process by automating the job application process. Utilizing artificial intelligence, it enables users to apply for multiple jobs in a tailored way.
★ 30kco-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5khawkeye. A perception system for ball tracking in cricket and tennis.
★ 54stauthlib. Demo Repo
★ 34fastapi-sso. FastAPI plugin to enable SSO to most common providers (such as Facebook login, Google login and login via Microsoft Office 365 Account)
★ 484EasyOCR. Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
★ 30kGOT-OCR2.0. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
★ 8.2kvideos. Code for the manim-generated scenes used in 3blue1brown videos
★ 11khuggingface_hub. The official CLI and Python client for the Hugging Face Hub.
★ 3.8kInstantMesh. InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
★ 4.5kvideohash. Near Duplicate Video Detection (Perceptual Video Hashing) - Get a 64-bit comparable hash-value for any video.
★ 378docling. Get your documents ready for gen AI
★ 64kDiffusionKit. On-device Image Generation for Apple Silicon
★ 703AISbod.
★ 1TAPTR. [ECCV 2024 & NeurIPS 2024 & ICLR 2026] Official implementation of the paper TAPTR & TAPTRv2 & TAPTRv3
★ 280Pillow. Python Imaging Library (fork)
★ 14kpython-progressbar. Progressbar 2 - A progress bar for Python 2 and Python 3 - "pip install progressbar2"
★ 879sports. computer vision and sports
★ 5.3ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kmasa. Official Implementation of CVPR24 highlight paper: Matching Anything by Segmenting Anything
★ 1.4kARC-AGI. The Abstraction and Reasoning Corpus
★ 4.8kmanim. Animation engine for explanatory math videos
★ 89kMAC-SQL. MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL
★ 344siamese-triplet. Siamese and triplet networks with online pair/triplet mining in PyTorch
★ 3.2knotebooks. A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge models like RF-DETR, YOLO11, SAM 3, and Qwen3-VL.
★ 9.6kreview_object_detection_metrics. Object Detection Metrics. 14 object detection metrics: mean Average Precision (mAP), Average Recall (AR), Spatio-Temporal Tube Average Precision (STT-AP). This project supports different bounding box formats as in COCO, PASCAL, Imagenet, etc.
★ 1.2kViTPose. The official repo for [NeurIPS'22] "ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation" and [TPAMI'23] "ViTPose++: Vision Transformer for Generic Body Pose Estimation"
★ 2.1kInstaGen. InstaGen: Enhancing Object Detection by Training on Synthetic Dataset, CVPR2024
★ 91omnimotion. Python
★ 2.3kpips. Particle Video Revisited
★ 603PerceptualSimilarity. LPIPS metric. pip install lpips
★ 4.3koptimum-benchmark. 🏋️ A unified multi-backend utility for benchmarking Transformers, Timm, PEFT, Diffusers and Sentence-Transformers with full support of Optimum's hardware optimizations & quantization schemes.
★ 338MotionAGFormer. Official implementation of the paper "MotionAGFormer: Enhancing 3D Pose Estimation with a Transformer-GCNFormer Network" (WACV 2024).
★ 232mediapipe-python-sample. MediaPipeのPythonパッケージのサンプルです。2024/9/1時点でPython実装のある15機能について用意しています。
★ 338vanna-streamlit. Vanna AI Streamlit App
★ 332sn-gamestate. [CVPRW'24] SoccerNet Game State Reconstruction: End-to-End Athlete Tracking and Identification on a Minimap (CVPR24 - CVSports workshop)
★ 430ScoreHMR. ScoreHMR: Score-Guided Diffusion for 3D Human Recovery (CVPR 2024)
★ 435lerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kvanna-slack. Slack bot for Vanna AI
★ 40vanna. 🤖 Chat with your SQL database 📊. Accurate Text-to-SQL Generation via LLMs using Agentic Retrieval 🔄.
★ 24kDiffMOT. code for CVPR2024 paper: DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction
★ 448gpt4all. GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77kllama_cloud_services. Knowledge Agents and Management in the Cloud
★ 4.3kpykan. Kolmogorov Arnold Networks
★ 16ksql-metadata. Uses tokenized query returned by python-sqlparse and generates query metadata
★ 880WrenAI. GenBI (Generative BI) for AI agents, an open-source, governed text-to-SQL through an open context layer that turns natural-language questions into trusted dashboards, charts, and SQL across 20+ data sources, such as BigQuery, Snowflake, PostgreSQL, ClickHouse, Amazon Redshift, Databricks and more.
★ 17kPOPE. Welcome to the project repository for POPE (Promptable Pose Estimation), a state-of-the-art technique for 6-DoF pose estimation of any object in any scene using a single reference.
★ 171openpose. OpenPose: Real-time multi-person keypoint detection library for body, face, hands, and foot estimation
★ 34kimagen-pytorch. Implementation of Imagen, Google's Text-to-Image Neural Network, in Pytorch
★ 8.4kcookbook. Open-source AI cookbook
★ 2.7kAutoScraper. Official implement of paper "AutoScraper: A Progressive Understanding Web Agent for Web Scraper Generation" [EMNLP 24']
★ 489ollama-python. Ollama Python library
★ 10kollama. Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
★ 177khatchet. 🪓 An orchestration engine for background tasks, AI agents, and durable workflows
★ 7.6kdino-tracker. Official Pytorch Implementation for “DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video” (ECCV 2024)
★ 566llama.cpp. LLM inference in C/C++
★ 122kSpaTracker. [CVPR 2024 Highlight] Official PyTorch implementation of SpatialTracker: Tracking Any 2D Pixels in 3D Space
★ 1.1k