This is your work, valued
class-balanced-loss. Class-Balanced Loss Based on Effective Number of Samples. CVPR 2019
★ 616cvpr18-inaturalist-transfer. Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning. CVPR 2018
★ 196cvpr18-caption-eval. Learning to Evaluate Image Captioning. CVPR 2018
★ 85dataset-granularity.
★ 12personal_website. Yin Cui's personal website
★ 2gmn. Generative Matching Networks
★ 1decaf-release. Decaf is DEPRECATED! Please visit http://caffe.berkeleyvision.org/ for Caffe, the new framework that has all the good things: GPU computation, full train/test scripts, native C++, and an active community!
★ 1mixup-cifar10. mixup: Beyond Empirical Risk Minimization
★ 1OmniShotCut. OmniShotCut is a sensitive and more informative SoTA on Shot Boundary Detection task.
★ 264learn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73kcosmos-reason2. Cosmos-Reason2 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
★ 432cosmos-rl. Cosmos-RL is a flexible and scalable Reinforcement Learning framework specialized for Physical AI applications.
★ 468sage. Official Code Release of SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
★ 374Awesome-World-Models. A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
★ 3.2kcosmos-predict2.5. Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the form of video.
★ 1.3kphysical-ai-bench. [CVPR 2026 Oral] PAI-Bench: A Comprehensive Benchmark for Physical AI
★ 91AgiBot-World. [IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
★ 3.1kNFT. Implementation of Negative-aware Finetuning (NFT) algorithm for "Bridging Supervised Learning and Reinforcement Learning in Math Reasoning"
★ 88gemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kcosmos-predict2. Cosmos-Predict2 is a collection of general-purpose world foundation models for Physical AI that can be fine-tuned into customized world models for downstream applications.
★ 794cosmos-xenna. Python library for building and running distributed data pipelines using Ray
★ 79cosmos-predict1. Cosmos-Predict1 is a collection of general-purpose world foundation models for Physical AI that can be fine-tuned into customized world models for downstream applications.
★ 465DaTaSeg-Objects365-Instance-Segmentation. We release the DaTaSeg Objects365 Instance Segmentation Dataset introduced in the DaTaSeg paper, which can be used as an evaluation benchmark for weakly or semi supervised segmentation.
★ 22describe-anything. [ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning
★ 1.5kcosmos-reason1. Cosmos-Reason1 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
★ 952cosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kTokenBench. A Video Tokenizer Evaluation Dataset
★ 158Cosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kMegatron-Energon. Megatron's multi-modal data loader
★ 374LLM101n. LLM101n: Let's build a Storyteller
★ 38kyarn. YaRN: Efficient Context Window Extension of Large Language Models
★ 1.7kdclm. DataComp for Language Models
★ 1.5kbuild-nanogpt. Video+code lecture on building nanoGPT from scratch
★ 5.4kllama3. The official Meta Llama 3 GitHub site
★ 29ktorchtitan. A PyTorch native platform for training generative AI models
★ 5.6kBlender-Donut-Tutorial. :doughnut: Beginner Blender 4.0 Tutorial
★ 2Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kawesome-6d-object. Awesome work on object 6 DoF pose estimation
★ 888dm-haiku. JAX-based neural network library
★ 3.3kAwesome-Transformer-Attention. An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites
★ 5.1kGrounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kCVinW_Readings. A collection of papers on the topic of ``Computer Vision in the Wild (CVinW)''
★ 1.4kopen_clip. An open source implementation of CLIP.
★ 14kd2l-en. Interactive deep learning book with multi-framework code, math, and discussions. Adopted at 500 universities from 70 countries including Stanford, MIT, Harvard, and Cambridge.
★ 29kuncertainty-baselines. High-quality implementations of standard and SOTA methods on a variety of tasks.
★ 1.6kvision_transformer. Jupyter Notebook
★ 13kAwesome-Video-Datasets. Video datasets
★ 1.7kGSAM. PyTorch repository for ICLR 2022 paper (GSAM) which improves generalization (e.g. +3.8% top-1 accuracy on ImageNet with ViT-B/32)
★ 147jonbarron.github.io. HTML
★ 3.6kdmvr. Python
★ 68newt. Natural World Tasks
★ 45TF_JAX_tutorials. All about the fundamental blocks of TF and JAX!
★ 280Awesome-Knowledge-Distillation. Awesome Knowledge-Distillation. 分类整理的知识蒸馏paper(2014-2021)。
★ 2.7kTransformer-in-Vision. Recent Transformer-based CV and related works.
★ 1.3kmdetr. Python
★ 1.1kAwesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 3.6kspatio-temporal-contrastive-film. Unsupervised Film Genre Classification using Spatio-Temporal Contrastive Learning
★ 32copy-paste-aug. Copy-paste augmentation for segmentation and detection tasks
★ 570CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kawesome-self-supervised-learning. A curated list of awesome self-supervised methods
★ 6.4kawesome-contrastive-self-supervised-learning. A comprehensive list of awesome contrastive self-supervised learning papers.
★ 1.3kWSL-Images. Weakly Supervised Learning On Images
★ 602jax. Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
★ 36ksimclr-converter. A PyTorch converter for SimCLR checkpoints
★ 108SlowFast. PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
★ 7.4kmoco. PyTorch implementation of MoCo: https://arxiv.org/abs/1911.05722
★ 5.1kSpineNet-Pytorch. SpineNet - mmdetection (Pytorch) Implementation
★ 68swav. PyTorch implementation of SwAV https//arxiv.org/abs/2006.09882
★ 2.1ktide. A General Toolbox for Identifying Object Detection Errors
★ 741Awesome-of-Long-Tailed-Recognition. A curated list of long-tailed recognition resources.
★ 580VTHCL. [arXiv 2020] Video Representation Learning with Visual Tempo Consistency
★ 24video-nonlocal-net. Non-local Neural Networks for Video Classification
★ 2kfashionpedia.
★ 173fashionpedia-api. Python API for Fashionpedia Dataset
★ 188noisystudent. Code for Noisy Student Training. https://arxiv.org/abs/1911.04252
★ 765SpineNet-Pytorch. This project is a kind of implementation of SpineNet(CVPR 2020) using mmdetection.
★ 99fucking-algorithm. Crack LeetCode, not only how, but also why.
★ 135kPSIS. Data Augmentation for Object Detection via Progressive and Selective Instance-Switching
★ 77scene-representation-networks. Official Pytorch implementation of Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations
★ 440pytorch-metric-learning. The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in PyTorch.
★ 6.3ktorchcv. TorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision
★ 2.3kCutMix-PyTorch. Official Pytorch implementation of CutMix regularizer
★ 1.3kdetectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35kLDAM-DRW. [NeurIPS 2019] Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
★ 702EfficientNet-PyTorch. A PyTorch implementation of EfficientNet
★ 8.2kBag_of_Tricks_for_Image_Classification_with_Convolutional_Neural_Networks. experiments on Paper <Bag of Tricks for Image Classification with Convolutional Neural Networks> and other useful tricks to improve CNN acc
★ 741Class-balanced-loss-pytorch. Pytorch implementation of the paper "Class-Balanced Loss Based on Effective Number of Samples"
★ 803pytorch-cifar. 95.47% on CIFAR10 with PyTorch
★ 6.4kmixup-cifar10. mixup: Beyond Empirical Risk Minimization
★ 1.2kpytorch_compact_bilinear_pooling. Compact Bilinear Pooling for PyTorch
★ 254librosa. Python library for audio and music analysis
★ 8.5kwaymo-open-dataset. Waymo Open Dataset
★ 3.4ksonnet. TensorFlow-based neural network library
★ 9.9klvis-api. Python API for LVIS Dataset
★ 430PointFlow. PointFlow : 3D Point Cloud Generation with Continuous Normalizing Flows
★ 867pyrobot. PyRobot: An Open Source Robotics Research Platform
★ 2.3kxlnet. XLNet: Generalized Autoregressive Pretraining for Language Understanding
★ 6.2kimat_fashion_comp. The iMaterialist Fashion Attribute Dataset
★ 89DeepFashion2. DeepFashion2 Dataset https://arxiv.org/pdf/1901.07973.pdf
★ 2.6kmodanet. ModaNet: A large-scale street fashion dataset with polygon annotations
★ 362kaggle-imaterialist. The First Place Solution of Kaggle iMaterialist (Fashion) 2019 at FGVC6
★ 484Made-With-ML. Learn how to develop, deploy and iterate on production-grade ML applications.
★ 49kmujoco-py. MuJoCo is a physics engine for detailed, efficient rigid body simulations with contacts. mujoco-py allows using MuJoCo from Python 3.
★ 3.1kimet-fgvcx.
★ 13SPADE. Semantic Image Synthesis with SPADE
★ 7.7kcvpr18-caption-eval. Learning to Evaluate Image Captioning. CVPR 2018
★ 85class-balanced-loss. Class-Balanced Loss Based on Effective Number of Samples. CVPR 2019
★ 616cvpr18-inaturalist-transfer. Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning. CVPR 2018
★ 196transferlearning. Transfer learning / domain adaptation / domain generalization / multi-task learning etc. Papers, codes, datasets, applications, tutorials.-迁移学习
★ 14kPlotNeuralNet. Latex code for making neural networks diagrams
★ 25kdataset-distillation. Open-source code for paper "Dataset Distillation"
★ 827GNNPapers. Must-read papers on graph neural networks (GNN)
★ 17kgraph_nets. Build Graph Nets in Tensorflow
★ 5.4kdeepcluster. Deep Clustering for Unsupervised Learning of Visual Features
★ 1.7kpytext. A natural language modeling framework based on PyTorch
★ 6.3kbert. TensorFlow code and pre-trained models for BERT
★ 40kmaskrcnn-benchmark. Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch.
★ 9.4kvid2vid. Pytorch implementation of our method for high-resolution (e.g. 2048x1024) photorealistic video-to-video translation.
★ 8.7kmmdetection. OpenMMLab Detection Toolbox and Benchmark
★ 33kmmcv. OpenMMLab Computer Vision Foundation
★ 6.5ksnca.pytorch. Improving Generalization via Scalable Neighborhood Component Analysis
★ 138TransferLearningClassification. Xuhong Li, Yves Grandvalet, and Franck Davoine. "Explicit Inductive Bias for Transfer Learning with Convolutional Networks." In ICML 2018.
★ 57pytorch-classification. Classification with PyTorch.
★ 1.7kconvnet-aig. PyTorch implementation for Convolutional Networks with Adaptive Inference Graphs
★ 184daily-paper-computer-vision. 记录每天整理的计算机视觉/深度学习/机器学习相关方向的论文
★ 6.8kcodalab-competitions. CodaLab Competitions
★ 539CatPapers. Cool vision, learning, and graphics papers on Cats!
★ 1.2kdopamine. Dopamine is a research framework for fast prototyping of reinforcement learning algorithms.
★ 11kself-critical.pytorch. Unofficial pytorch implementation for Self-critical Sequence Training for Image Captioning. and others.
★ 1k3d-recon. Implementation for paper "Learning Single-View 3D Reconstruction with Limited Pose Supervision".
★ 64mmf. A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
★ 5.6kpytorch-maml-rl. Reinforcement Learning with Model-Agnostic Meta-Learning in Pytorch
★ 884unbiased-offline-recommender-evaluation. Jupyter Notebook
★ 35talkthewalk. This repository provides code for reproducing experiments of the paper Talk The Walk: Navigating New York City Through Grounded Dialogue by Harm de Vries, Kurt Shuster, Dhruv Batra, Devi Parikh, Jason Weston, and Douwe Kiela.
★ 112DPNs. Dual Path Networks
★ 526intrinsic-dimension. Jupyter Notebook
★ 219lemniscate.pytorch. Unsupervised Feature Learning via Non-parametric Instance Discrimination
★ 758cmr. Project repo for Learning Category-Specific Mesh Reconstruction from Image Collections
★ 485explain_teach. Teaching Categories to Human Learners with Visual Explanations - CVPR 2018
★ 11Relation-Networks-for-Object-Detection. Relation Networks for Object Detection
★ 1.1kSuperPointPretrainedNetwork. PyTorch pre-trained model for real-time interest point detection, description, and sparse tracking (https://arxiv.org/abs/1712.07629)
★ 2.2kwechat_friends. 微信好友信息分析并可视化以及自动回复微信消息
★ 1.3khub. A library for transfer learning by reusing parts of TensorFlow models.
★ 3.5klow-shot-shrink-hallucinate. Presenting Low-shot Visual Recognition by Shrinking and Hallucinating Features
★ 310CSrankings. A web app for ranking computer science departments according to their research output in selective venues, and for finding active faculty across a wide range of areas.
★ 3.2kSelective-Joint-Fine-tuning. Codes and models for the CVPR 2017 spotlight paper "Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-tuning".
★ 79seg_every_thing. Code release for Hu et al., Learning to Segment Every Thing. in CVPR, 2018.
★ 423CompactBilinearPooling-Pytorch. A Pytorch Implementation for Compact Bilinear Pooling.
★ 187Paper_Reading_List. Recommended Papers. Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Learning (cs.LG)
★ 713incubator-mxnet. Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler; for Python, R, Julia, Scala, Go, Javascript and more
★ 114ELF. ELF: a platform for game research with AlphaGoZero/AlphaZero reimplementation
★ 3.4kFoodx.
★ 17fgvcx_flower_comp.
★ 18readme2tex. Renders TeXy Math for Github Readme - No longer needed with official MathTex support on GH
★ 911fgvcx_fungi_comp. FGVCx Fungi Classfiication competition details
★ 53MUNIT. Multimodal Unsupervised Image-to-Image Translation
★ 2.7kminigo. An open-source implementation of the AlphaGoZero algorithm
★ 3.5kmxnet. Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler; for Python, R, Julia, Scala, Go, Javascript and more
★ 21kava-dataset. The AVA dataset densely annotates 80 atomic visual actions in 351k movie clips with actions localized in space and time, resulting in 1.65M action labels with multiple labels per human occurring frequently.
★ 348CollMetric. A Tensorflow implementation of Collaborative Metric Learning (CML)
★ 158influence-release. Jupyter Notebook
★ 809inat_comp_2018. CNN training code for iNaturalist 2018 image classification competition
★ 85tpu. Reference models and tools for Cloud TPUs.
★ 5.3kstate-of-the-art-result-for-machine-learning-problems. This repository provides state of the art (SoTA) results for all machine learning problems. We do our best to keep this repository up to date. If you do find a problem's SoTA result is out of date or missing, please raise this as an issue or submit Google form (with this information: research paper name, dataset, metric, source code and year). We will fix it immediately.
★ 8.9kjulia. The Julia Programming Language
★ 49kopenrec. OpenRec is an open-source and modular library for neural network-inspired recommendation algorithms
★ 417Neural-Style-Transfer-Papers. :pencil2: Neural Style Transfer: A Review
★ 1.6kDeformable-ConvNets. Deformable Convolutional Networks
★ 4.1kMeta-Learning-Papers. Meta Learning / Learning to Learn / One Shot Learning / Few Shot Learning
★ 2.7kDetectron. FAIR's research platform for object detection research, implementing popular algorithms like Mask R-CNN and RetinaNet.
★ 26krelative_depth. Code for the NIPS 2016 paper
★ 123places-coco2017.github.io. Places + COCO Challenge
★ 3coco-analyze. A wrapper of the COCOeval class for extended keypoint error estimation analysis.
★ 232Mask_RCNN. Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow
★ 26kpyemd. [Deprecated; use POT instead] Fast EMD for Python: a wrapper for Pele and Werman's C++ implementation of the Earth Mover's Distance metric
★ 492handong1587.github.io. CSS
★ 3.1kneural-editor. Repository for "Generating Sentences by Editing Prototypes"
★ 332gmn. Generative Matching Networks
★ 29cocodataset.github.io. HTML
★ 286SMASH. An experimental technique for efficiently exploring neural architectures.
★ 491SENet. Squeeze-and-Excitation Networks
★ 3.6ktensorflow_compact_bilinear_pooling. Compact Bilinear Pooling in TensorFlow
★ 140DiscoGAN. Official implementation of "Learning to Discover Cross-Domain Relations with Generative Adversarial Networks"
★ 777the-gan-zoo. A list of all named GANs!
★ 15kmodels. Models and examples built with TensorFlow
★ 78ktfrecords. Functions for creating tfrecords for TensorFlow models.
★ 110tf_classification. Training, evaluation and testing code for image classification using TensorFlow
★ 134imat_comp. iMaterialist (Fashion) Competition 2019
★ 164gym. A toolkit for developing and comparing reinforcement learning algorithms.
★ 37kuniverse. Universe: a software platform for measuring and training an AI's general intelligence across the world's supply of games, websites and other applications.
★ 7.5kbaselines. OpenAI Baselines: high-quality implementations of reinforcement learning algorithms
★ 17kpytorch_fft. PyTorch wrapper for FFTs
★ 317inat_comp_2017. CNN training code for iNaturalist 2017 image classification competition
★ 17cocostuff10k. The official homepage of the (outdated) COCO-Stuff 10K dataset.
★ 280Hard-Aware-Deeply-Cascaed-Embedding. source code for the paper "Hard-Aware-Deeply-Cascaed-Embedding"
★ 32fairseq-lua. Facebook AI Research Sequence-to-Sequence Toolkit
★ 3.7ksignal. Signal Processing Library for PyTorch
★ 39pytorch-CycleGAN-and-pix2pix. Image-to-Image Translation in PyTorch
★ 25kinat_comp. iNaturalist competition details
★ 813hugo. The world’s fastest framework for building websites.
★ 89kkit. 🧱 Describe your site, AI builds it, you own it as Markdown. Snap together Tailwind blocks like Lego — landing pages, blogs, portfolios, docs & more. No AI slop. Free to deploy anywhere 👇
★ 9.6ktriplet-network-pytorch. Python
★ 378visdom. A flexible tool for creating, organizing, and sharing visualizations of live, rich data. Supports Torch and Numpy https://visdom.dev
★ 10kawesome-deep-learning-papers. The most cited deep learning papers
★ 26kshow-attend-and-tell. TensorFlow Implementation of "Show, Attend and Tell"
★ 906ImageCaptioning.pytorch. I decide to sync up this repo and self-critical.pytorch. (The old master is in old master branch for archive)
★ 1.5kthe-incredible-pytorch. The Incredible PyTorch: a curated list of tutorials, papers, projects, communities and more relating to PyTorch.
★ 13k