This is your work, valued
Cheers!
Dataset-REPAIR. REPresentAtion bIas Removal (REPAIR) of datasets
★ 57valhalla-nmt. Code repository for CVPR 2022 paper "VALHALLA: Visual Hallucination for Machine Translation"
★ 28svitt. Code for CVPR 2023 paper "SViTT: Temporal Learning of Sparse Video-Text Transformers"
★ 21bg-resample-ood. Background resampling for out-of-distribution detection
★ 13oh-my-codex. OmX - Oh My codeX: Your codex is not alone. Add hooks, agent teams, HUDs, and so much more.
★ 32kAttention-Residuals.
★ 3.4ktinyworlds. A minimal implementation of DeepMind's Genie world model
★ 1.3kreasoning-from-scratch. Implement a reasoning LLM in PyTorch from scratch, step by step
★ 4.8kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kgemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kSpatialLM. [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
★ 4.6kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kDeepSeek-VL2. DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
★ 5.3khymba. Python
★ 214granite-snack-cookbook. Granite Snack Cookbook -- easily consumable recipes (python notebooks) that showcase the capabilities of the Granite models
★ 387Cosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7kpapers_we_read. Summaries for exciting works in the field of Deep Learning.
★ 357ChatDev. ChatDev 2.0: Dev All through LLM-powered Multi-Agent Collaboration
★ 34kLLaVA-NeXT. Python
★ 4.7kEgoVLPv2. Code release for "EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone" [ICCV, 2023]
★ 110gpt_paper_assistant. GPT4 based personalized ArXiv paper assistant bot
★ 547syn-rep-learn. Learning from synthetic data - code and models
★ 328grok-1. Grok open release
★ 52kSegment-Everything-Everywhere-All-At-Once. [NeurIPS 2023] Official implementation of the paper "Segment Everything Everywhere All at Once"
★ 4.8kNExT-GPT. Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6kDallEval. DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models (ICCV 2023)
★ 143al-folio. A beautiful, simple, clean, and responsive Jekyll theme for academics
★ 16kAcademic-project-page-template. A project page template for academic papers. Demo at https://eliahuhorwitz.github.io/Academic-project-page-template/
★ 5.1ktime-diffusion. Official code repo for "Editing Implicit Assumptions in Text-to-Image Diffusion Models"
★ 89visprog. Official code for VisProg (CVPR 2023 Best Paper!)
★ 774ControlNet-v1-1-nightly. Nightly release of ControlNet 1.1
★ 5.2kdebias-gensynth. Balancing the Picture: Debiasing Vision-Language Datasets with Synthetic Contrast Sets
★ 12LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kTPT. Test-time Prompt Tuning (TPT) for zero-shot generalization in vision-language models (NeurIPS 2022))
★ 214textgen. Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.
★ 48kControlNet. Let us control diffusion models!
★ 34koptimum-habana. Easy and lightning fast training of 🤗 Transformers on Habana Gaudi processor (HPU)
★ 212debias_vl. Code for Debiasing Vision-Language Models via Biased Prompts
★ 60ChatGPT. Reverse engineered ChatGPT API
★ 28kinstruct-pix2pix. Python
★ 6.9kCrisscrossed-Captions. Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO
★ 54multimodal. TorchMultimodal is a PyTorch library for training state-of-the-art multimodal multi-task models at scale.
★ 1.7kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kLM_bias. [ICML 2021] Towards Understanding and Mitigating Social Biases in Language Models
★ 61nanoGPT. The simplest, fastest repository for training/finetuning medium-sized GPTs.
★ 62kNetworks-Beyond-Attention. A compilation of network architectures for vision and others without usage of self-attention mechanism
★ 81Awesome-Diffusion-Models. A collection of resources and papers on Diffusion Models
★ 12kMeMViT. Code Release for MeMViT Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition, CVPR 2022
★ 155mmengine. OpenMMLab Foundational Library for Training Deep Learning Models
★ 1.5kkinetics-dataset. Shell
★ 982awesome-tips.
★ 4.7kDeltaCNN. DeltaCNN End-to-End CNN Inference of Sparse Frame Differences in Videos
★ 59mvit. Code Release for MViTv2 on Image Recognition.
★ 456ConvNeXt. Code release for ConvNeXt model
★ 6.4kyt-dlp. A feature-rich command-line audio/video downloader
★ 181kSparseConvNet. Submanifold sparse convolutional networks
★ 2.1kgdown. Google Drive public file downloader when curl/wget fails.
★ 5.3kcomposer. Supercharge Your Model Training
★ 5.5kMTTR. Python
★ 655math. The MATH Dataset (NeurIPS 2021)
★ 1.4kvisualbert. Code for the paper "VisualBERT: A Simple and Performant Baseline for Vision and Language"
★ 542Revisit-MMT. Python
★ 25mdetr. Python
★ 1.1kDALLE-pytorch. Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
★ 5.6kCLIP-ViL. [ICLR 2022] code for "How Much Can CLIP Benefit Vision-and-Language Tasks?" https://arxiv.org/abs/2107.06383
★ 419wit. WIT (Wikipedia-based Image Text) Dataset is a large multimodal multilingual dataset comprising 37M+ image-text sets with 11M+ unique images across 100+ languages.
★ 1.1kflores. Facebook Low Resource (FLoRes) MT Benchmark
★ 771CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kminiClip. Python
★ 48CondensedMovies. Story-Based Retrieval with Contextual Embeddings. Largest freely available movie video dataset. [ACCV'20]
★ 205XLM. PyTorch original implementation of Cross-lingual Language Model Pretraining.
★ 2.9kpytorchvideo. A deep learning library for video understanding research.
★ 3.6kfairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
★ 32kpytorch-lightning. Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
★ 31ksmart_open. Utils for streaming large files (S3, HDFS, gzip, bz2...)
★ 3.5kVMZ. VMZ: Model Zoo for Video Modeling
★ 1.1kSlowFast. PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
★ 7.4kmmaction. An open-source toolbox for action understanding based on PyTorch
★ 1.9kar-cutpaste. Cut and paste your surroundings using AR
★ 15kawesome. 😎 Awesome lists about all kinds of interesting topics
★ 490kCATER. CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
★ 112temporal-shift-module. [ICCV 2019] TSM: Temporal Shift Module for Efficient Video Understanding
★ 2.2kdetectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35knatural-adv-examples. A Harder ImageNet Test Set (CVPR 2021)
★ 621Optical-Flow-GPU-Docker. Compute dense optical flow using TV-L1 algorithm with NVIDIA GPU acceleration.
★ 58PyVideoResearch. A repository of common methods, datasets, and tasks for video research
★ 536CVPR2019. Displays all the 2019 CVPR Accepted Papers in a way that they are easy to parse.
★ 73nvvl. A library that uses hardware acceleration to load sequences of video frames to facilitate machine learning training
★ 693PyAV. Pythonic bindings for FFmpeg's libraries.
★ 3.2kfootball. Check out the new game server:
★ 3.7ktexture-vs-shape. Pre-trained models, data, code & materials from the paper "ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness" (ICLR 2019 Oral)
★ 812Domain-generalization. All about domain generalization
★ 845AlphaPose. Real-Time and Accurate Full-Body Multi-Person Pose Estimation&Tracking System
★ 8.6kDetectAndTrack. The implementation of an algorithm presented in the CVPR18 paper: "Detect-and-Track: Efficient Pose Estimation in Videos"
★ 998awesome-action-recognition. A curated list of action recognition and related area resources
★ 4kAction_Recognition_Zoo. Codes for popular action recognition models, verified on the something-something data set.
★ 246dense_flow. OpenCV Implementation of different optical flow algorithms
★ 235pytorch-rotationnet. Python
★ 82temporal-fields. Code for training temporal fully-connected CRF models in Torch
★ 69Conference-Acceptance-Rate. Acceptance rates for the major AI conferences
★ 4.8kmanim. Animation engine for explanatory math videos
★ 89kNon-local_pytorch. Implementation of Non-local Block.
★ 1.6ktowards-realistic-predictors. pytorch implementation for paper, towards realistic predictors
★ 17Detectron.pytorch. A pytorch implementation of Detectron. Both training from scratch and inferring directly from pretrained Detectron weights are available.
★ 2.8ktwo-stream-action-recognition. Using two stream architecture to implement a classic action recognition method on UCF101 dataset
★ 874kinetics_i3d_pytorch. Inflated i3d network with inception backbone, weights transfered from tensorflow
★ 5473D-ResNets-PyTorch. 3D ResNets for Action Recognition (CVPR 2018)
★ 4kTorchFusion. A modern deep learning framework built to accelerate research and development of AI systems
★ 258skorch. A scikit-learn compatible neural network library that wraps PyTorch
★ 6.2kfaster-rcnn.pytorch. A faster pytorch implementation of faster r-cnn
★ 7.9ktnt. A lightweight library for PyTorch training tools and utilities
★ 1.7kpytorch-custom-dataset-examples. Some custom dataset examples for PyTorch
★ 874pretrained-models.pytorch. Pretrained ConvNets for pytorch: NASNet, ResNeXt, ResNet, InceptionV4, InceptionResnetV2, Xception, DPN, etc.
★ 9.1ktensorboardX. tensorboard for pytorch (and chainer, mxnet, numpy, ...)
★ 8kpytorch-playground. Base pretrained models and datasets in pytorch (MNIST, SVHN, CIFAR10, CIFAR100, STL10, AlexNet, VGG16, VGG19, ResNet, Inception, SqueezeNet)
★ 2.7kvisdom. A flexible tool for creating, organizing, and sharing visualizations of live, rich data. Supports Torch and Numpy https://visdom.dev
★ 10kpytorch-tutorial. PyTorch Tutorial for Deep Learning Researchers
★ 32kpytorch-meta-optimizer. A PyTorch implementation of Learning to learn by gradient descent by gradient descent
★ 315pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
★ 102k