This is your work, valued
Research Scientist, MIT CSAIL
ast. Code for the Interspeech 2021 paper "AST: Audio Spectrogram Transformer".
★ 1.5kltu. Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
★ 478ssast. Code for the AAAI 2022 paper "SSAST: Self-Supervised Audio Spectrogram Transformer".
★ 430whisper-at. Code and Pretrained Models for Interspeech 2023 Paper "Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong Audio Event Taggers"
★ 422cav-mae. Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".
★ 292gopt. Code for the ICASSP 2022 paper "Transformer-Based Multi-Aspect Multi-Granularity Non-native English Speaker Pronunciation Assessment".
★ 218vocalsound. Dataset and baseline code for the VocalSound dataset (ICASSP2022).
★ 165psla. Code for the TASLP paper "PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation".
★ 150uavm. Code for the IEEE Signal Processing Letters 2022 paper "UAVM: Towards Unifying Audio and Visual Models".
★ 57python-compute-eer. Simple Python script to compute equal error rate (EER) for machine learning model evaluation.
★ 41ReMASC. ReMASC: Realistic Replay Attack Corpus for Voice Controlled Systems
★ 38llm_speech_emotion_challenge. Jupyter Notebook
★ 23realtime-adversarial-attack. Code for IJCAI 2019 paper "Real-time Adversarial Attack".
★ 20multichannel-antispoof. Code for SPL paper "Detecting Replay Attacks Using Multi-Channel Audio: A Neural Network-Based Method"
★ 7awesome-whisper. 🔊 Awesome list for Whisper — an open-source AI-powered speech recognition system developed by OpenAI
★ 5efficient-voice-antispoof. Jupyter Notebook
★ 4Awesome-Multimodal-Large-Language-Models. Latest Papers and Datasets on Multimodal Large Language Models
★ 4ESC-50. ESC-50: Dataset for Environmental Sound Classification
★ 3Best-Audio-Classification-Resources-with-Deep-learning. List of articles related to deep learning applied to music
★ 3kaldi-abbr. kaldi name convention note
★ 2Facebook-Interview-Coding-1. 个人整理的Facebook实习面试题目解法,时间范围2016.8-2017.3
★ 2yuangongnd.github.io. HTML
★ 2SincNet. SincNet is a neural architecture for efficiently processing raw audio samples.
★ 1Speech_DB_Engine. Python
★ 1whisper-flamingo. Whisper-Flamingo [Interspeech 2024] and mWhisper-Flamingo [IEEE SPL 2025] for Audio-Visual Speech Recognition and Translation
★ 210llm_speech_emotion_challenge. Jupyter Notebook
★ 23SAIL. SAIL: Search Augmented Instruction Learning
★ 160cav-mae. Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".
★ 292whisper-at. Code and Pretrained Models for Interspeech 2023 Paper "Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong Audio Event Taggers"
★ 422ltu. Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
★ 478uavm. Code for the IEEE Signal Processing Letters 2022 paper "UAVM: Towards Unifying Audio and Visual Models".
★ 57awesome-self-supervised-learning. A curated list of awesome self-supervised methods
★ 6.4kpytorch-randaugment. Unofficial PyTorch Reimplementation of RandAugment.
★ 635whisper. Robust Speech Recognition via Large-Scale Weak Supervision
★ 106kpaper-reading. 深度学习经典、新论文逐段精读
★ 34kawesome-grounding. awesome grounding: A curated list of research papers in visual grounding
★ 1.1klatent-diffusion. High-Resolution Image Synthesis with Latent Diffusion Models
★ 14kdataset. Multi30k Dataset
★ 192google-scholar-crawler. Crawl google scholar with least code.
★ 46crawler-google-scholar. This bot crawls and downloads statistics and pictures from google scholar's researchers.
★ 21gscholar-citations-crawler. Crawl all your citations from Google Scholar
★ 58kaldi-diar-latte. steps to perform text-based speaker diarization with kaldi toolkit
★ 12faiss. A library for efficient similarity search and clustering of dense vectors.
★ 41khowto100m. Code for the HowTo100M paper
★ 304everything_at_once. Official implementation of "Everything at Once - Multi-modal Fusion Transformer for Video Retrieval." CVPR 2022
★ 115voice_datasets. 🔊 A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).
★ 2.2kdmvr. Python
★ 68vocalsound. Dataset and baseline code for the VocalSound dataset (ICASSP2022).
★ 165awesome-video-feature-extractor. Video feature extractor in PyTorch.
★ 8VGGSound. VGGSound: A Large-scale Audio-Visual Dataset
★ 359awesome-audio-visual. A curated list of different papers and datasets in various areas of audio-visual processing
★ 775liwc-python. Linguistic Inquiry and Word Count (LIWC) analyzer
★ 240video_feature_extractor. Easy to use video deep features extractor
★ 322AVLnet. Code for the AVLnet (Interspeech 2021) and Cascaded Multilingual (Interspeech 2021) papers.
★ 54scenic. Scenic: A Jax Library for Computer Vision Research and Beyond
★ 3.8kbyol-pytorch. Usable Implementation of "Bootstrap Your Own Latent" self-supervised learning, from Deepmind, in Pytorch
★ 1.9kmerlot_reserve. Code release for "MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound"
★ 146MLIC-KD-WSD. Multi-Label Image Classification via Knowledge Distillation from Weakly-Supervised Detection (ACM MM 2018)
★ 59ctc-segmentation. Segment an audio file and obtain utterance alignments. (Python package)
★ 348espnet. ESPnet with Lightweight Sinc Convolutions and CTC segmentation patches
★ 5pase. Problem Agnostic Speech Encoder
★ 446ssast. Code for the AAAI 2022 paper "SSAST: Self-Supervised Audio Spectrogram Transformer".
★ 430lottery-ticket-hypothesis. A reimplementation of "The Lottery Ticket Hypothesis" (Frankle and Carbin) on MNIST.
★ 731Lottery-Ticket-Hypothesis-in-Pytorch. This repository contains a Pytorch implementation of the paper "The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks" by Jonathan Frankle and Michael Carbin that can be easily adapted to any model/dataset.
★ 349deconstructing-lottery-tickets. Python
★ 143PyTorch_CIFAR10. Pretrained TorchVision models on CIFAR10 dataset (with weights)
★ 692open_lth. A repository in preparation for open-sourcing lottery ticket hypothesis code.
★ 640A-journey-into-Convolutional-Neural-Network-visualization-. A journey into Convolutional Neural Network visualization
★ 275BERT-like-is-All-You-Need. The code for our INTERSPEECH 2020 paper - Jointly Fine-Tuning "BERT-like'" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
★ 121svcca. Jupyter Notebook
★ 644PretrainCNNwithNoData. Python
★ 3Audio-Classification. Pytorch code for "Rethinking CNN Models for Audio Classification"
★ 129psla. Code for the TASLP paper "PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation".
★ 150gopt. Code for the ICASSP 2022 paper "Transformer-Based Multi-Aspect Multi-Granularity Non-native English Speaker Pronunciation Assessment".
★ 218pytorch-ssd. MobileNetV1, MobileNetV2, VGG based SSD/SSD-lite implementation in Pytorch 1.0 / Pytorch 0.4. Out-of-box support for retraining on Open Images dataset. ONNX and Caffe2 support. Experiment Ideas like CoordConv.
★ 1.4kpytorch-cnn-visualizations. Pytorch implementation of convolutional neural network visualization techniques
★ 8.2kfcn.berkeleyvision.org. Fully Convolutional Networks for Semantic Segmentation by Jonathan Long*, Evan Shelhamer*, and Trevor Darrell. CVPR 2015 and PAMI 2016.
★ 3.4kunderstanding-transfer-learning. Python
★ 45deeplab2. DeepLab2 is a TensorFlow library for deep labeling, aiming to provide a unified and state-of-the-art TensorFlow codebase for dense pixel labeling tasks.
★ 1kimagenet-sample-images. 1000 images, one per image-net class. For easy visualization/exploration of classes.
★ 185perceiver-pytorch. Implementation of Perceiver, General Perception with Iterative Attention, in Pytorch
★ 1.2kUnsupSeg. Self-Supervised Contrastive Learning for Unsupervised Phoneme Segmentation (INTERSPEECH 2020)
★ 147ShapeTextureDebiasedTraining. Code and models for the paper Shape-Texture Debiased Neural Network Training (ICLR 2021)
★ 112augmentation-corruption. This repository provides code for "On Interaction Between Augmentations and Corruptions in Natural Corruption Robustness".
★ 47local-global-features-cnn. Code for my Master's Thesis: "The Role of Local Versus Global Features in Convolutional Neural Networks"
★ 2model-vs-human. Benchmark your model on out-of-distribution datasets with carefully collected human comparison data (NeurIPS 2021 Oral)
★ 362texture-vs-shape. Pre-trained models, data, code & materials from the paper "ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness" (ICLR 2019 Oral)
★ 812Stylized-ImageNet. Code to create Stylized-ImageNet, a stylized version of standard ImageNet (ICLR 2019 Oral)
★ 528torchcrepe. Pytorch implementation of the CREPE pitch tracker
★ 523imbalanced-regression. [ICML 2021, Long Talk] Delving into Deep Imbalanced Regression
★ 916transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kkaldi-io-for-python. Python functions for reading kaldi data formats. Useful for rapid prototyping with python.
★ 378kaldiio. A pure python module for reading and writing kaldi ark files
★ 268voxpopuli. A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
★ 574libri-light. dataset for lightly supervised training using the librivox audio book recordings. https://librivox.org/.
★ 528ML-zoo. Python
★ 263video-contrastive-learning. Video Contrastive Learning with Global Context, ICCVW 2021
★ 162yt-search. JavaScript
★ 121BERT-flow. TensorFlow implementation of On the Sentence Embeddings from Pre-trained Language Models (EMNLP 2020)
★ 535captum. Model interpretability and understanding for PyTorch
★ 5.7kunilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kSpeechSplit. Unsupervised Speech Decomposition Via Triple Information Bottleneck
★ 697youtube-dl. Command-line program to download videos from YouTube.com and other video sites
★ 141ks3prl. Self-Supervised Speech Pre-training and Representation Learning Toolkit
★ 2.6klinear-attention-transformer. Transformer based on a variant of attention that is linear complexity in respect to sequence length
★ 842linformer. Implementation of Linformer for Pytorch
★ 307sam. SAM: Sharpness-Aware Minimization (PyTorch)
★ 2kknowledge-distillation-pytorch. A PyTorch implementation for exploring deep and shallow knowledge distillation (KD) experiments with flexibility
★ 2kKnowledge-Distillation-Zoo. Pytorch implementation of various Knowledge Distillation (KD) methods.
★ 1.8kInterspeech2021. This repository describes our reproducible framework for assessing self-supervised representation learning from speech
★ 52label-errors. 🛠️ Corrected Test Sets for ImageNet, MNIST, CIFAR, Caltech-256, QuickDraw, IMDB, Amazon Reviews, 20News, and AudioSet
★ 188cleanlab. Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
★ 12kNoise-Contrastive-Estimation-NCE-for-pyTorch. Re-implementation of the Noise Contrastive Estimation algorithm for pyTorch, following "Noise-contrastive estimation: A new estimation principle for unnormalized statistical models." (Gutmann and Hyvarinen, AISTATS 2010)
★ 45CPC_audio. An implementation of the Contrast Predictive Coding (CPC) method to train audio features in an unsupervised fashion.
★ 374awesome-knowledge-distillation. Awesome Knowledge Distillation
★ 3.9kknowledge-distillation-papers. knowledge distillation papers
★ 765Awesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 3.6kInformer2020. The GitHub repository for the paper "Informer" accepted by AAAI 2021.
★ 6.5kT2T-ViT. ICCV2021, Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
★ 1.2kbertviz. BertViz: Visualize Attention in Transformer Models
★ 8.1kespnet. End-to-End Speech Processing Toolkit
★ 9.9kaxial-attention. Implementation of Axial attention - attending to multi-dimensional data efficiently
★ 396pytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37kViT-pytorch. Pytorch reimplementation of the Vision Transformer (An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale)
★ 2.2keinops. Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)
★ 9.6kx-transformers. A concise but complete full-attention transformer with a set of promising experimental features from various papers
★ 5.9kvit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25kbaidudl. This is a multi-thread download tool for pan.baidu.com
★ 205awesome-kaldi. This is a list of features, scripts, blogs and resources for better using Kaldi ( http://kaldi-asr.org/ )
★ 536PlotNeuralNet. Latex code for making neural networks diagrams
★ 25kml-visuals. 🎨 ML Visuals contains figures and templates which you can reuse and customize to improve your scientific writing.
★ 17kvoxceleb-ivector. Voxceleb1 i-vector based speaker recognition system
★ 43Speaker-Recognition. This repo contains my attempt to create a Speaker Recognition and Verification system using SideKit-1.3.1
★ 116Awesome-Learning-with-Label-Noise. A curated list of resources for Learning with Noisy Labels
★ 2.7kkaldi_lm. Old language modeling tool that's used in kaldi
★ 17imbalanced-dataset-sampler. A (PyTorch) imbalanced dataset sampler for oversampling low frequent classes and undersampling high frequent ones.
★ 2.3kgender-recognition-by-voice. Identify a voice as male or female.
★ 33opensmile-python. Python package for openSMILE
★ 329surfboard. Novoic's audio feature extraction library
★ 437attention-is-all-you-need-pytorch. A PyTorch implementation of the Transformer model in "Attention is All You Need".
★ 9.8kbc_learning_sound. Chainer implementation of between-class learning for sound recognition https://arxiv.org/abs/1711.10282
★ 95mixup. Implementation of the mixup training method
★ 469audioset_tagging_cnn. Python
★ 1.8kpytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
★ 102kEfficientNet-PyTorch. A PyTorch implementation of EfficientNet
★ 8.2kskynet-ddp-slurm-example. Example of using PyTorch DistributedDataParallel and SLURM on skynet
★ 31distributed_tutorial. Python
★ 263google-research. Google Research
★ 38kapex. A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
★ 9kfairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
★ 32kContrastive-Predictive-Coding-PyTorch. Contrastive Predictive Coding for Automatic Speaker Verification
★ 506Best-Audio-Classification-Resources-with-Deep-learning. List of articles related to deep learning applied to music
★ 102forced-alignment-tools. A collection of links and notes on forced alignment tools
★ 942Autoregressive-Predictive-Coding. Autoregressive Predictive Coding: An unsupervised autoregressive model for speech representation learning
★ 191Attentive-Filtering-Network. University of Edinbrugh-Johns Hopkins University's system for ASVspoof 2017 Version 2.0 dataset.
★ 50ontology. The Audio Set Ontology aims to provide a comprehensive set of categories to describe sound events.
★ 714COVID-19-train-audio. COVID-19 Coughs files for training AI models
★ 42kaldi-abbr. kaldi name convention note
★ 2multichannel-antispoof. Code for SPL paper "Detecting Replay Attacks Using Multi-Channel Audio: A Neural Network-Based Method"
★ 7efficient-voice-antispoof. Jupyter Notebook
★ 4kaldi. kaldi-asr/kaldi is the official location of the Kaldi project.
★ 15ksampleCNN-pytorch. Pytorch implementation of "Sample-level Deep Convolutional Neural Networks for Music Auto-tagging Using Raw Waveforms"
★ 1examples. A set of examples around pytorch in Vision, Text, Reinforcement Learning, etc.
★ 24kSincNet. SincNet is a neural architecture for efficiently processing raw audio samples.
★ 1.2klhvqt. Frontend filterbank learning module with HVQT initialization capabilities.
★ 21deep_complex_networks. Implementation related to the Deep Complex Networks
★ 789wav2letter. Facebook AI Research's Automatic Speech Recognition Toolkit
★ 6.4kpytorch-gradual-warmup-lr. Gradually-Warmup Learning Rate Scheduler for PyTorch
★ 987python-compute-eer. Simple Python script to compute equal error rate (EER) for machine learning model evaluation.
★ 41pyroomacoustics. Pyroomacoustics is a package for audio signal processing for indoor applications. It was developed as a fast prototyping platform for beamforming algorithms in indoor scenarios.
★ 1.9kmic_array. DOA, VAD and KWS for ReSpeaker Microphone Array
★ 332Facebook-Interview-Coding-1. 个人整理的Facebook实习面试题目解法,时间范围2016.8-2017.3
★ 18BERT-BiLSTM-CRF-NER. Tensorflow solution of NER task Using BiLSTM-CRF model with Google BERT Fine-tuning And private Server services
★ 4.9kName-Entity-Recognition. Lstm-crf,Lattice-CRF,bert-ner及近年ner相关论文follow
★ 567ND_replay-attack. Dataset and codes of the audio replay attack research in ND Mobile Computing Lab
★ 3sequence_tagging. Named Entity Recognition (LSTM + CRF) - Tensorflow
★ 2ktf_ner. Simple and Efficient Tensorflow implementations of NER models with tf.estimator and tf.data
★ 924CRF. Linear chain conditional random fields are implemented using Numpy and Mxnet/Gluon, and batch training is supported, not limited to training the data one by one.
★ 22realtime-adversarial-attack. Code for IJCAI 2019 paper "Real-time Adversarial Attack".
★ 20bert. TensorFlow code and pre-trained models for BERT
★ 40kasvspoof2017. an implement of asvspoof 2017 using pytorch
★ 21ReMASC_Exp. Baseline Experiments for ReMASC dataset.
★ 6DeepSpeech. DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
★ 27kReMASC. ReMASC: Realistic Replay Attack Corpus for Voice Controlled Systems
★ 38Decoupled-Learning-Conditional-GAN. Decoupled Learning for Conditional Adversarial Networks
★ 17pix2pix. Image-to-image translation with conditional adversarial nets
★ 11kwavegan. WaveGAN: Learn to synthesize raw audio with generative adversarial networks
★ 1.4krpgan. RP-GAN: Stable GAN Training with Random Projections
★ 22paper-learning-to-pivot. Repository for the paper "Learning to Pivot with Adversarial Networks"
★ 35dcgan_tensorflow. Tensorflow implementation of "UNSUPERVISED REPRESENTATION LEARNING WITH DEEP CONVOLUTIONAL GENERATIVE ADVERSARIAL NETWORKS"
★ 81text-to-image. Generative Adversarial Text to Image Synthesis / Please Star -->
★ 599gan. A 1D toy example of optimizing a generative model using the WGAN-GP model.
★ 25pyAudioAnalysis. Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications
★ 6.3kword2vec-api. Simple web service providing a word embedding model
★ 1.4kcleverhans. An adversarial example library for constructing attacks, building defenses, and benchmarking both
★ 6.5kadversarial_BFGS_TensorFlow. Adversarial example creation based on BFGS algorithm implemented under TensorFlow
★ 6