This is your work, valued
MGMatting. This repository includes the official project of Mask Guided (MG) Matting, presented in our paper: Mask Guided Matting via Progressive Refinement Network
★ 374GG-Transformer. Code and models for the paper Glance-and-Gaze Vision Transformer
★ 28CAKES. This repository contains the code for our AAAI2021 paper CAKES: Channel-wise Automatic KErnel Shrinking for Efficient 3D Networks.
★ 12Switchable-Normalization. Code for Switchable Normalization from "Differentiable Learning-to-Normalize via Switchable Normalization", https://arxiv.org/abs/1806.10779
★ 1video2dataset. Easily create large video dataset from video urls
★ 1webdataset. A high-performance Python-based I/O system for large (and small) deep learning problems, with strong support for PyTorch.
★ 1flops-counter.pytorch. Flops counter for convolutional networks in pytorch framework
★ 1vision. Datasets, Transforms and Models specific to Computer Vision
★ 1ponytail. Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
★ 92kKimi-K3. Open Frontier Intelligence
★ 7.5kao. PyTorch native quantization and sparsity for training and inference
★ 2.9krllab. rllab is a framework for developing and evaluating reinforcement learning algorithms, fully compatible with OpenAI Gym.
★ 3.1ki1. Code release for "i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models"
★ 255DeepEP. DeepEP: an efficient expert-parallel communication library
★ 9.9kFreqFlow. The official implementation of "Frequency-Aware Flow Matching for High-Quality Image Generation"
★ 29deltatok. [CVPR 2026 Highlight] A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
★ 217BAR. [ICML 2026] code & model for arxiv paper "Autoregressive Image Generation with Masked Bit Modeling"
★ 60openpi. Python
★ 13kEmu3.5. Native Multimodal Models are World Learners
★ 1.5kScale-RAE. Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
★ 255VTP. [ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generation
★ 496JiT. PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
★ 2.5kFlowTok. PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokens
★ 17lightly-studio. Curate, Annotate, and Manage Your Data in LightlyStudio.
★ 871RAE. Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
★ 2kdatatrove. Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.
★ 3.2kdinov3. Reference PyTorch implementation and models for DINOv3
★ 11kunitree_ros. C++
★ 1.5klivox_ros_driver2. Livox device driver under Ros(Compatible with ros and ros2), support Lidar HAP and Mid-360.
★ 820Vision-Language-Vision. Python
★ 65NaVILA. [RSS'25] This repository is the implementation of "NaVILA: Legged Robot Vision-Language-Action Model for Navigation"
★ 676token-opt. Code for ICML 2025 Paper "Highly Compressed Tokenizer Can Generate Without Training"
★ 206SelftokTokenizer. Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
★ 238GRAT. This repository includes the official implementation of our paper "Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers"
★ 56xAR. This repository includes the official implementation of our paper "Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation"
★ 251Diffusion-wo-CFG. Official Implementation for Diffusion Models Without Classifier-free Guidance
★ 175fractalgen. PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
★ 1.2kGFT. Python
★ 531.58bit.flux.
★ 279LARP. Official Pytorch implementation for LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior (ICLR 2025 Oral).
★ 107Infinity. [CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
★ 1.6kmaskbit. Implementation of the paper "MaskBit: Embedding-free Image Generation from Bit Tokens"
★ 94ViCaS. ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation (CVPR'25)
★ 21CrossFlow. [CVPR2025] PyTorch-based reimplementation of CrossFlow, as proposed in 'Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution'
★ 345FlowAR. “FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching” FlowAR employs a simplest scale design and is compatible with any VAE.
★ 171genex. Generative World Explorer
★ 167adaptive-length-tokenizer. Adaptive Length Image Tokenization via Recurrent Allocation | How many tokens is an image worth ?
★ 146CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kmar. PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838
★ 1.9kOmniTokenizer. [NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.
★ 325Megatron-LM. Ongoing research training transformer models at scale
★ 17ktitok-pytorch. Implementation of TiTok, proposed by Bytedance in "An Image is Worth 32 Tokens for Reconstruction and Generation"
★ 1841d-tokenizer. This repo contains the code for 1D tokenizer and generator
★ 1.2kDiMR. [NeurIPS 24] Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models
★ 44MambaOut. MambaOut: Do We Really Need Mamba for Vision? (CVPR 2025)
★ 2.7kvideo2dataset. Easily create large video dataset from video urls
★ 662CVQ-VAE. [ICCV 2023] Online Clustered Codebook
★ 189content-debiased-fvd. [CVPR 2024] On the Content Bias in Fréchet Video Distance
★ 148coconut_cvpr2024. Jupyter Notebook
★ 204ViTamin. [CVPR 2024] Official implementation of "ViTamin: Designing Scalable Vision Models in the Vision-language Era"
★ 211VideoGPT. Jupyter Notebook
★ 1.1kmistral-inference. Official inference library for Mistral models
★ 11kopen-muse. Open reproduction of MUSE for fast text2image generation.
★ 358LVM. Python
★ 1.8kAxial-VS. This repo contains the code for our TMLR paper: A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
★ 27mmc4. MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.
★ 954qa-lora. Official PyTorch implementation of QA-LoRA
★ 147Emu. Emu Series: Generative Multimodal Models from BAAI
★ 1.8kOmniScient-Model. This repo contains the code for our paper Towards Open-Ended Visual Recognition with Large Language Model
★ 102CLIPSelf. [ICLR2024 Spotlight] Code Release of CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
★ 207DeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43k3D-TransUNet. This is the official repository for the paper "3D TransUNet: Advancing Medical Image Segmentation through Vision Transformers"
★ 315llama_index. LlamaIndex is the leading document agent and OCR platform
★ 51kDETA. Detection Transformers with Assignment
★ 270gpt4free. The official gpt4free repository | various collection of powerful language models | opus 4.6 gpt 5.3 kimi 2.5 deepseek v3.2 gemini 3
★ 67kV3Det. Python
★ 121shikra. Python
★ 814lvis-api. Python API for LVIS Dataset
★ 430fc-clip. [NeurIPS 2023] This repo contains the code for our paper Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP
★ 345all-seeing. [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of the Open World"
★ 506GPT4RoI. (ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
★ 555open_flamingo. An open-source framework for training large multimodal models.
★ 4.1kkmax-deeplab. a PyTorch re-implementation of ECCV 2022 paper based on Detectron2: k-means mask Transformer.
★ 80LAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kllama. Inference code for Llama models
★ 60kEDICT. Jupyter Notebook
★ 322Detic. Code release for "Detecting Twenty-thousand Classes using Image-level Supervision".
★ 2kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kprompt-to-prompt. Jupyter Notebook
★ 3.5kCrossAttentionControl. Unofficial implementation of "Prompt-to-Prompt Image Editing with Cross Attention Control" with Stable Diffusion
★ 1.3kLaCLIP. [NeurIPS 2023] Text data, code and pre-trained models for paper "Improving CLIP Training with Language Rewrites"
★ 291SAN. Open-vocabulary Semantic Segmentation
★ 384MOSS. An open-source tool-augmented conversational language model from Fudan University
★ 12kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kSegment-Everything-Everywhere-All-At-Once. [NeurIPS 2023] Official implementation of the paper "Segment Everything Everywhere All at Once"
★ 4.8kprompt-pretraining. Official implementation for the paper "Prompt Pre-Training with Over Twenty-Thousand Classes for Open-Vocabulary Visual Recognition"
★ 259Grounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kPainter. Painter & SegGPT Series: Vision Foundation Models from BAAI
★ 2.6kUNINEXT. [CVPR'23] Universal Instance Perception as Object Discovery and Retrieval
★ 1.3kCRIS.pytorch. An official PyTorch implementation of the CRIS paper
★ 282LoRA. Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"
★ 14ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kZegFormer. Official code for "Decoupling Zero-Shot Semantic Segmentation"
★ 180CAT-Seg. Official Implementation of "CAT-Seg🐱: Cost Aggregation for Open-Vocabulary Semantic Segmentation"
★ 388MaskCLIP. Official PyTorch implementation of "Extract Free Dense Labels from CLIP" (ECCV 22 Oral)
★ 481img2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4.4kDiT. Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
★ 8.7kControlNet. Let us control diffusion models!
★ 34kODISE. Official PyTorch implementation of ODISE: Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models [CVPR 2023 Highlight]
★ 946GroupViT. Official PyTorch implementation of GroupViT: Semantic Segmentation Emerges from Text Supervision, CVPR 2022.
★ 788CutLER. Code release for "Cut and Learn for Unsupervised Object Detection and Instance Segmentation" and "VideoCutLER: Surprisingly Simple Unsupervised Video Instance Segmentation"
★ 1.1kEVA. EVA Series: Visual Representation Fantasies from BAAI
★ 2.7kopen_clip. An open source implementation of CLIP.
★ 14kPartImageNet. Introduction and scripts for the paper "PartImageNet: A Large, High-Quality Dataset of Parts" (Ju He, Shuo Yang, Shaokang Yang, Adam Kortylewski, Xiaoding Yuan, Jie-Neng Chen, Shuai Liu, Cheng Yang, Alan Yuille).
★ 137H-Deformable-DETR. [CVPR2023] This is an official implementation of paper "DETRs with Hybrid Matching".
★ 279parti.
★ 1.6kautogluon. Fast and Accurate ML in 3 Lines of Code
★ 11kdeeplab2. DeepLab2 is a TensorFlow library for deep labeling, aiming to provide a unified and state-of-the-art TensorFlow codebase for dense pixel labeling tasks.
★ 1kdxvk. Vulkan-based implementation of D3D8, 9, 10 and 11 for Linux / Wine
★ 18kSwin-Transformer-TF. Tensorflow implementation of Swin Transformer model.
★ 216ModelsGenesis. [MICCAI 2019 Young Scientist Award] [MEDIA 2020 Best Paper Award] Models Genesis, one of the first "foundation" models in medical image analysis for multiple downstream tasks
★ 789L2B. This repository includes the official project of L2B, from our paper "Learning to Bootstrap for Combating Label Noise".
★ 32SAGE. No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models (ICLR 2022)
★ 29convnext-tf. Python
★ 30sceneparsing. Development kit for MIT Scene Parsing Benchmark
★ 470fast_advprop. [ICLR 2022]: Fast AdvProp
★ 35LabelAssemble. [ISBI 2023] Official Implementation for Label-Assemble
★ 20ConvNeXt. Code release for ConvNeXt model
★ 6.4kmae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4kMask2Former. Code release for "Masked-attention Mask Transformer for Universal Image Segmentation"
★ 3.4kTransMix. [CVPR 2022] This repository includes the official project for the paper: TransMix: Attend to Mix for Vision Transformers.
★ 158ViTs-vs-CNNs. [NeurIPS 2021]: Are Transformers More Robust Than CNNs? (Pytorch implementation & checkpoints)
★ 180K-Net. [NeurIPS2021] Code Release of K-Net: Towards Unified Image Segmentation
★ 484CaliCO. Code for ICCV2021 paper: Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real Images
★ 15MaskFormer. Per-Pixel Classification is Not All You Need for Semantic Segmentation (NeurIPS 2021, spotlight)
★ 1.5kInExtremIS. Official PyTorch implementation of InExtremIS
★ 22GG-Transformer. Code and models for the paper Glance-and-Gaze Vision Transformer
★ 28relabel_imagenet. Python
★ 406Swin-Transformer. This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
★ 16kPanopticFCN. Fully Convolutional Networks for Panoptic Segmentation (CVPR2021 Oral)
★ 404detectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35knoah-research. Noah Research
★ 973PixPro. Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning, CVPR 2021
★ 364transformer-in-transformer. Implementation of Transformer in Transformer, pixel level attention paired with patch level attention for image classification, in Pytorch
★ 306IIC. Invariant Information Clustering for Unsupervised Image Classification and Segmentation
★ 891PVT. Official implementation of PVT series
★ 1.9kpytorch-lightning. Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
★ 31kirn. Weakly Supervised Learning of Instance Segmentation with Inter-pixel Relations, CVPR 2019 (Oral)
★ 535Unsupervised-Semantic-Segmentation. Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals. [ICCV 2021]
★ 413pytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37kvissl. VISSL is FAIR's library of extensible, modular and scalable components for SOTA Self-Supervised Learning with images.
★ 3.3kTransUNet. This repository includes the official project of TransUNet, presented in our paper: TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation.
★ 3.2kpytorch-cnn-visualizations. Pytorch implementation of convolutional neural network visualization techniques
★ 8.2knerf. Code release for NeRF (Neural Radiance Fields)
★ 11kT2T-ViT. ICCV2021, Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
★ 1.2kmmselfsup. OpenMMLab Self-Supervised Learning Toolbox and Benchmark
★ 3.3kdeit. Official DeiT repository
★ 4.4kself-label. Self-labelling via simultaneous clustering and representation learning. (ICLR 2020)
★ 546CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kViT-pytorch. Pytorch reimplementation of the Vision Transformer (An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale)
★ 2.2kvision-transformer-pytorch. Pytorch version of Vision Transformer (ViT) with pretrained models. This is part of CASL (https://casl-project.github.io/) and ASYML project.
★ 365DALLE-pytorch. Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
★ 5.6kmmsegmentation. OpenMMLab Semantic Segmentation Toolbox and Benchmark.
★ 9.9kSETR. [CVPR 2021 & IJCV 2024] Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
★ 1.1kHearthstone-Deck-Tracker. A deck tracker and deck manager for Hearthstone on Windows
★ 4.9kbyol-pytorch. Usable Implementation of "Bootstrap Your Own Latent" self-supervised learning, from Deepmind, in Pytorch
★ 1.9kMOS_Meticulous-Object-Segmentation.
★ 57swav. PyTorch implementation of SwAV https//arxiv.org/abs/2006.09882
★ 2.1kMGMatting. This repository includes the official project of Mask Guided (MG) Matting, presented in our paper: Mask Guided Matting via Progressive Refinement Network
★ 374Adversarial_Metric_Attack. The code of "Adversarial Metric Attack for Person Re-identification"
★ 30CAKES. This repository contains the code for our AAAI2021 paper CAKES: Channel-wise Automatic KErnel Shrinking for Efficient 3D Networks.
★ 12ViP-DeepLab. Python
★ 227BNET. Batch Normalization with Enhanced Linear Transformation
★ 53AdelaiDet. AdelaiDet is an open source toolbox for multiple instance-level detection and recognition tasks.
★ 3.5kvit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25kpixel-level-contrastive-learning. Implementation of Pixel-level Contrastive Learning, proposed in the paper "Propagate Yourself", in Pytorch
★ 266MODNet. A Trimap-Free Portrait Matting Solution in Real Time [AAAI 2022]
★ 4.3kdetr. End-to-End Object Detection with Transformers
★ 15kclosed-form-matting. Python implementation of A. Levin D. Lischinski and Y. Weiss. A Closed Form Solution to Natural Image Matting. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), June 2006, New York
★ 444GFM. [IJCV 2022] Bridging Composite and Real: Towards End-to-end Deep Image Matting
★ 940ShapeTextureDebiasedTraining. Code and models for the paper Shape-Texture Debiased Neural Network Training (ICLR 2021)
★ 112senet.pytorch. PyTorch implementation of SENet
★ 2.3kCascadePSP. [CVPR 2020] CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local Refinement
★ 886SynthCP. Offical code base for the ECCV oral paper "Synthesize then Compare: Detecting Failures and Anomalies for Semantic Segmentation"
★ 60self-supervised-3d-tasks. Python
★ 192unnas. Code for "Are labels necessary for neural architecture search"
★ 93PyContrast. PyTorch implementation of Contrastive Learning methods
★ 2kCAM. Class Activation Mapping
★ 1.9kmmagic. OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
★ 7.4kCnC_Remastered_Collection. Command & Conquer: Remastered Collection
★ 21kFeatureLearningRotNet. Python
★ 526RAdam. On the Variance of the Adaptive Learning Rate and Beyond
★ 2.5kDetectoRS. DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution
★ 1.1kPatchAttack. PatchAttack (ECCV 2020)
★ 65FBA_Matting. Official repository for the paper F, B, Alpha Matting
★ 502GCA-Matting. Official repository for Natural Image Matting via Guided Contextual Attention
★ 406Background-Matting. Background Matting: The World is Your Green Screen
★ 4.8kipyvolume. 3d plotting for Python in the Jupyter notebook based on IPython widgets using WebGL
★ 2kindexnet_matting. (ICCV'19) Indices Matter: Learning to Index for Deep Image Matting
★ 396AutoNL. Code for AutoNL on ImageNet (CVPR2020)
★ 104mmaction. An open-source toolbox for action understanding based on PyTorch
★ 1.9k3D-ResNets-PyTorch. 3D ResNets for Action Recognition (CVPR 2018)
★ 4kpytorch-classification. Classification with PyTorch.
★ 1.7kAtomNAS. [ICLR 2020]: 'AtomNAS: Fine-Grained End-to-End Neural Architecture Search'
★ 220temporal-segment-networks. Code & Models for Temporal Segment Networks (TSN) in ECCV 2016
★ 1.6krethinking-network-pruning. Rethinking the Value of Network Pruning (Pytorch) (ICLR 2019)
★ 1.5ksingle-path-nas. Single-Path NAS: Designing Hardware-Efficient ConvNets in less than 4 Hours
★ 394