This is your work, valued
Contrastive-Predictive-Coding-PyTorch. Contrastive Predictive Coding for Automatic Speaker Verification
★ 506pytorch-kaldi-neural-speaker-embeddings. A light weight neural speaker embeddings extraction based on Kaldi and PyTorch.
★ 136ASSERT. JHU's system submission to the ASVspoof 2019 Challenge: Anti-Spoofing with Squeeze-Excitation and Residual neTworks (ASSERT).
★ 57Attentive-Filtering-Network. University of Edinbrugh-Johns Hopkins University's system for ASVspoof 2017 Version 2.0 dataset.
★ 50Unsupervised-TTS. Python
★ 42Semi-Supervsied-Spoken-Language-Understanding-PyTorch. Semi-supervised spoken language understanding (SLU) via self-supervised speech and language model pretraining
★ 12PARP-wav2vec-PyTorch. Python
★ 9LSTM. Voice activity detection of noisy speech files with LSTM. LSTM is implemented with Keras. Data processing is done with Python, MATLAB, and Bash. Experiments are done on Johns Hopkins CLSP GPUs.
★ 6scale. Some of my public work at https://hltcoe.jhu.edu/research/scale/scale-2017/
★ 3mfcc. Steps to extract mfcc features from audios
★ 2intro-machine-learning-paper. Paper I have read for understanding statistical machine learning and speaker recognition
★ 2noise_cancellation. Noise cancellation system for speech files with different kinds of noise.
★ 2Intro_Mechatronics. Introduction to Mechatronics (EN.520.240) is an introductory robotic course I took in Fall 2016 at Johns Hopkins University. The course consists of 7 labs and a final project, and each one builds up to the next one. The codes are written for Arduino, and each subfile corresponds to a lab.
★ 1LDA. An implementation of Linear Discriminant Analysis on multi-class speaker recognition.
★ 1VGNSL. [ACL 2019] Visually Grounded Neural Syntax Acquisition
★ 1lexicon-learner. Python
★ 1TTS-Pruning-Pytorch. Python
★ 1MultiTalk. [NeurIPS 2025] Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
★ 3kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kxformers. Hackable and optimized Transformers building blocks, supporting a composable construction.
★ 11kaxlearn. An Extensible Deep Learning Library
★ 2.4kseamless_communication. Foundational Models for State-of-the-Art Speech and Text Translation
★ 12ktorchtune. PyTorch native post-training library
★ 5.8kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kxDiT. xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
★ 2.7kcodex. Lightweight coding agent that runs in your terminal
★ 103kBagel. Open-source unified multimodal model
★ 6.1kmoshi. Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
★ 11kLlama-AVSR. Official Pytorch implementation of "Large Language Models are Strong Audio-Visual Speech Recognition Learners" [ICASSP 2025] and "Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs" [ICASSP 2026].
★ 64BigCodec. Official implementation of the paper "BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec"
★ 218stable-audio-tools. Generative models for conditional audio generation
★ 3.8kMovieGenBench. Movie Gen Bench - two media generation evaluation benchmarks released with Meta Movie Gen
★ 442Mamba-ASR. ConMamba for Automatic Speech Recognition
★ 106mar. PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838
★ 1.9kSALMONN. SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5kDenseAV. Offical code for the CVPR 2024 Paper: Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
★ 88syllabify. Python module for syllabifying English ARPABET transcriptions
★ 74Amphion. Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
★ 10kRepCodec. Models and code for RepCodec: A Speech Representation Codec for Speech Tokenization
★ 196SpeechTokenizer. This is the code for the SpeechTokenizer presented in the SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models. Samples are presented on
★ 658GroupViT. Official PyTorch implementation of GroupViT: Semantic Segmentation Emerges from Text Supervision, CVPR 2022.
★ 788vq-vae-2-pytorch. Implementation of Generating Diverse High-Fidelity Images with VQ-VAE-2 in PyTorch
★ 1.8kLLaMA-Adapter. [ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
★ 5.9kss-phoneme-seg. Code for "Phoneme Segmentation Using Self-Supervised Speech Models", Strgar & Harwath, Proceedings of the IEEE Spoken Language Technology Workshop (SLT) 2023
★ 55alqalign. multilingual speech aligner
★ 78uavm. Code for the IEEE Signal Processing Letters 2022 paper "UAVM: Towards Unifying Audio and Visual Models".
★ 57vqwordseg. Unsupervised phone and word segmentation using dynamic programming on self-supervised VQ features.
★ 39awesome-neural-reprogramming-prompting. A curated list of awesome adversarial reprogramming and input prompting methods for neural networks since 2022
★ 40word-discovery. Word Discovery in Visually Grounded, Self-Supervised Speech Models
★ 27n-grammer-pytorch. Implementation of N-Grammer, augmenting Transformers with latent n-grams, in Pytorch
★ 81whisper. Robust Speech Recognition via Large-Scale Weak Supervision
★ 106kstructured-uncertainty. Python
★ 10gmm-hmm-asr. Python implementation of simple GMM and HMM models for isolated digit recognition.
★ 68xcfg. X (weighted / probabilistic) Context-Free Grammars
★ 25s2v_rc. Speech2Vec Reality Check
★ 88cpcfg. Fast and Modularized CFG-focused Models
★ 23rclone. "rsync for cloud storage" - Google Drive, S3, Dropbox, Backblaze B2, One Drive, Swift, Hubic, Wasabi, Google Cloud Storage, Azure Blob, Azure Files, Yandex Files
★ 59kgdrive. Google Drive CLI Client
★ 9kdiffwave-sr. Jupyter Notebook
★ 87parti.
★ 1.6kspoken_sent_embedding. Unsupervised spoken sentence embeddings
★ 14stable-diffusion. A latent text-to-image diffusion model
★ 73kS3-Router. [NeurIPS 2022] "Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing" by Yonggan Fu, Yang Zhang, Kaizhi Qian, Zhifan Ye, Zhongzhi Yu, Cheng-I Lai, Yingyan Lin
★ 17mega. Sequence modeling with Mega.
★ 303denoising-diffusion-model. A simple guide to diffusion models. Helpful in understanding the concept and practicing with the method.
★ 224cliora. Official codebase for ICLR oral paper Unsupervised Vision-Language Grammar Induction with Shared Structure Modeling
★ 36diffusion_models. A series of tutorial notebooks on denoising diffusion probabilistic models in PyTorch
★ 721diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kget-started-with-JAX. The purpose of this repo is to make it easy to get started with JAX, Flax, and Haiku. It contains my "Machine Learning with JAX" series of tutorials (YouTube videos and Jupyter Notebooks) as well as the content I found useful while learning about the JAX ecosystem.
★ 783contentvec. speech self-supervised representations
★ 520silero-vad. Silero VAD: pre-trained enterprise-grade Voice Activity Detector
★ 9.8ksilero-models. Silero Models: pre-trained text-to-speech models made embarrassingly simple
★ 6kDiffCloth. Code repository for our paper DiffCloth: Differentiable Cloth Simulation with Dry Frictional Contact
★ 431FaST-VGS-Family. Transformer-based visually grounded speech models
★ 19MAE-AST-Public. Public Code for the paper MAE-AST: Masked Autoencoding Audio Spectrogram Transformer
★ 93sinkhorn-simultrans. Implementation of the paper "Anticipation-Free Training for Simultaneous Machine Translation"
★ 8iTerm2-Color-Schemes. Over 450 terminal color schemes/themes for iTerm/iTerm2. Includes ports to Terminal, Konsole, PuTTY, Xresources, XRDB, Remmina, Termite, XFCE, Tilda, FreeBSD VT, Terminator, Kitty, MobaXterm, LXTerminal, Microsoft's Windows Terminal, Visual Studio, Alacritty, Ghostty, and many more
★ 27ktortoise-tts. A multi-voice TTS system trained with an emphasis on quality
★ 15kLASER. Language-Agnostic SEntence Representations
★ 3.7kdeep-rl-class. This repo contains the Hugging Face Deep Reinforcement Learning Course.
★ 5kicefall. Python
★ 1.5kmetaseq. Repo for external large-scale work
★ 6.6kspeech-resynthesis. An official reimplementation of the method described in the INTERSPEECH 2021 paper - Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.
★ 416DinkyTrain. Princeton NLP's pre-training library based on fairseq with DeepSpeed kernel integration 🚃
★ 117rnn-hierarchical-biases. Code for "Does syntax need to grow on trees? Sources of inductive bias in sequence to sequence networks"
★ 24DiffCSE. Code for the NAACL 2022 long paper "DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings"
★ 298dm_aux. Python
★ 67diora. Deep Inside-Outside Recursive Autoencoder
★ 89Paper-Writing-Tips. MLNLP社区用来帮助大家避免论文投稿小错误的整理仓库。 Paper Writing Tips
★ 4.6kawesome-align. A neural word aligner based on multilingual BERT
★ 379vgnsl_analysis. "What is Learned in Visually Grounded Neural Syntax Acquisition", Noriyuki Kojima, Hadar Averbuch-Elor, Alexander Rush and Yoav Artzi (ACL 2020)
★ 12VLGrammar. Python
★ 28CLMR. Official PyTorch implementation of Contrastive Learning of Musical Representations
★ 338recipes. Jupyter Notebook
★ 2wav2vec. a simplified version of wav2vec(1.0, vq, 2.0) in fairseq
★ 170tokenizations. Robust and Fast tokenizations alignment library for Rust and Python https://tamuhey.github.io/tokenizations/
★ 195mit-eecs-thesis-proposal-template. MIT EECS Thesis Proposal Template
★ 61pytorch-struct. Fast, general, and tested differentiable structured prediction in PyTorch
★ 1.1kself-attentive-parser. High-accuracy NLP parser with models for 11 languages.
★ 911unsupervised-parsing-tutorial. Unsupervised Natural Language Parsing (Tutorial)
★ 22Tree-Transformer. Implementation of the paper Tree Transformer
★ 219svgling. linguistics tree drawing to SVG in python, aimed at Jupyter
★ 64nltk. NLTK Source
★ 15kdenoiser. Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities.
★ 1.9ksurvey. A Survey on Neural Speech Synthesis https://arxiv.org/pdf/2106.15561.pdf
★ 371transformer-ls. Official PyTorch Implementation of Long-Short Transformer (NeurIPS 2021).
★ 228ffcv. FFCV: Fast Forward Computer Vision (and other ML workloads!)
★ 3kUniSpeech. UniSpeech - Large Scale Self-Supervised Learning for Speech
★ 486av_hubert. A self-supervised learning framework for audio-visual speech
★ 996s4. Structured state space sequence models
★ 2.9kslue-toolkit. A toolkit for Spoken Language Understanding Evaluation (SLUE) benchmark. Refer paper https://arxiv.org/abs/2111.10367 for more details. Official website: https://asappresearch.github.io/slue-toolkit/
★ 65sparseml. Libraries for applying sparsification recipes to neural networks with a few lines of code, enabling faster and smaller models
★ 2.1kvsepp. PyTorch Code for the paper "VSE++: Improving Visual-Semantic Embeddings with Hard Negatives"
★ 523AVLnet. Code for the AVLnet (Interspeech 2021) and Cascaded Multilingual (Interspeech 2021) papers.
★ 54Jacinle. Personal python toolbox.
★ 145Dropbox-Uploader. Dropbox Uploader is a BASH script which can be used to upload, download, list or delete files from Dropbox, an online file sharing, synchronization and backup service.
★ 6.6kunilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kminGPT. A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
★ 25kParallel-Tacotron2. PyTorch Implementation of Google's Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling
★ 191LossUpAccUp. Loss and accuracy go opposite ways...right?
★ 93oneshot. Shell
★ 3leetcode_solutions. Python
★ 13CPC_audio. An implementation of the Contrast Predictive Coding (CPC) method to train audio features in an unsupervised fashion.
★ 10Wav2Vec2_PyCTCDecode. Small repo describing how to use Hugging Face's Wav2Vec2 with PyCTCDecode
★ 110studyFiles. 一些电子书pdf
★ 314vpcfg. Visually Grounded PCFG Induction
★ 38lexical. Paper: Lexicon Learning for Few-Shot Neural Sequence Modeling
★ 17VGNSL. [ACL 2019] Visually Grounded Neural Syntax Acquisition
★ 90myprosody. A Python library for measuring the acoustic features of speech (simultaneous speech, high entropy) compared to ones of native speech.
★ 275textgrid. A Python module for interacting with Praat TextGrid files. Also includes a class for reading HTK .mlf files into Praat
★ 302librispeech-alignments. Word alignments generated by the Montreal Forced Aligner for the Librispeech dataset
★ 182Montreal-Forced-Aligner. Command line utility for forced alignment using Kaldi
★ 1.9kvoxceleb_trainer. In defence of metric learning for speaker recognition
★ 1.2kVGGVox. VGGVox models for Speaker Identification and Verification trained on the VoxCeleb (1 & 2) datasets
★ 402PaddleNLP. Easy-to-use and powerful LLM and SLM library with awesome model zoo.
★ 13ksmallfry. Python
★ 19voxpopuli. A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
★ 574torchdrug. A powerful and flexible machine learning platform for drug discovery
★ 1.6kdisentanglement_lib. disentanglement_lib is an open-source library for research on learning disentangled representations.
★ 1.4kpytorch-ivectors. GPU accelerated implementation of i-vector extractor training using PyTorch. Requires Kaldi for feature extraction and UBM training. An example script is provided for VoxCeleb data.
★ 63submitit. Python 3.8+ toolbox for submitting jobs to Slurm
★ 1.6kdeit. Official DeiT repository
★ 4.4kast. Code for the Interspeech 2021 paper "AST: Audio Spectrogram Transformer".
★ 1.5klong-range-arena. Long Range Arena for Benchmarking Efficient Transformers
★ 788lite-transformer. [ICLR 2020] Lite Transformer with Long-Short Range Attention
★ 609AutoPST. Global Rhythm Style Transfer Without Text Transcriptions
★ 285webMUSHRA. a MUSHRA compliant web audio API based experiment software
★ 10autovc. AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
★ 1.1kbyol-a. BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation
★ 237gdown. Google Drive public file downloader when curl/wget fails.
★ 5.3kvits. VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
★ 7.9kmcd. Mel cepstral distortion (MCD) computations in python.
★ 231hifi-gan. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
★ 2.4kespnet_model_zoo. ESPnet Model Zoo
★ 260sam. SAM: Sharpness-Aware Minimization (PyTorch)
★ 2kdatasets-CMU_Wilderness. CMU Wilderness Multilingual Speech Dataset
★ 292hyperfuture. Code for the paper Learning the Predictability of the Future (CVPR 2021)
★ 173Taylor_pruning. Pruning Neural Networks with Taylor criterion in Pytorch
★ 323invariant_rationalization. Tensorflow implementation of Invariant Rationalization
★ 50open_lth. A repository in preparation for open-sourcing lottery ticket hypothesis code.
★ 640CLIP4Clip. An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
★ 1kpythainlp. Thai natural language processing in Python
★ 1.1kshrinkbench. PyTorch library to facilitate development and standardized evaluation of neural network pruning methods.
★ 433GST-Tacotron. A PyTorch implementation of Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
★ 374GigaSpeech. Large, modern dataset for speech recognition
★ 731thesis. Danqi Chen's PhD Thesis
★ 227DALL-E. PyTorch package for the discrete VAE used for DALL·E.
★ 11kCLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kAwesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 3.6kBiLSTM_Collocation_Parser. BiLSTM+ELMo 搭建的中文 Collocation Parser
★ 9espnet-semi-supervised. ESPnet extensions for semi-supervised end-to-end speech recognition. See also https://github.com/ShigekiKarita/espnet-semi-supervised/tree/karita-asrtts for newer code in ICASSP2019 Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders
★ 38datasets. 🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
★ 22kdissect. Code for the Proceedings of the National Academy of Sciences 2020 article, "Understanding the Role of Individual Units in a Deep Neural Network"
★ 310comet. [ICLR 2021] Concept Learners for Few-Shot Learning
★ 115STT. 🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
★ 2.6kTTS-papers. 🐸 collection of TTS papers
★ 732TTS-recipes. 🐸TTS recipes for different datasets
★ 88TTS. 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 46kspeechbrain. A PyTorch-based Speech Toolkit
★ 12kuncertainty_benchmarking. Various code/notebooks to benchmark different ways we could estimate uncertainty in ML predictions.
★ 44uncertainty-toolbox. Uncertainty Toolbox: a Python toolbox for predictive uncertainty quantification, calibration, metrics, and visualization
★ 2kspeech-representations. Code for DeCoAR (ICASSP 2020) and BERTphone (Odyssey 2020)
★ 104maskbert. Python
★ 20VIBNet. Compressing Neural Networks using the Variational Information Bottleneck
★ 66tf-kaldi-speaker. Neural speaker recognition/verification system based on Kaldi and Tensorflow
★ 31DilatedRNN. Tensorflow implementation for DilatedRNN
★ 352self-supervised-speech-recognition. speech to text with self-supervised learning based on wav2vec 2.0 framework
★ 380yolov5. Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
★ 58kcv-dataset. Metadata and versioning details for the Common Voice dataset
★ 173diffwave. DiffWave is a fast, high-quality neural vocoder and waveform synthesizer.
★ 885light-weight-refinenet. Light-Weight RefineNet for Real-Time Semantic Segmentation
★ 740refinenet-pytorch. RefineNet-101 VOC in PyTorch
★ 197refinenet. RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
★ 605CPC_audio. An implementation of the Contrast Predictive Coding (CPC) method to train audio features in an unsupervised fashion.
★ 374powerai. This repo contains ancillary information used to assist users of IBM Watson Machine Learning Community Edition. This repo will contain How To's, Readme's, Dockerfiles, etc. that can be consumed by users looking to get started.
★ 57dscore. Diarization scoring tools.
★ 270Ax. Adaptive Experimentation Platform
★ 2.8kdenoising-diffusion-pytorch. Implementation of Denoising Diffusion Probabilistic Model in Pytorch
★ 11kGeneralizing-Lottery-Tickets. This repository contains code to replicate the experiments given in NeurIPS 2019 paper "One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers"
★ 50Lottery-ticket-hypothesis. This repository contains a Pytorch implementation of the article "The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks" and an application of this hypothesis to reinforcement learning
★ 9pytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37kpytext. A natural language modeling framework based on PyTorch
★ 6.3kdecaNLP. The Natural Language Decathlon: A Multitask Challenge for NLP
★ 2.3kDALLE-pytorch. Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
★ 5.6ksvoice. We provide a PyTorch implementation of the paper Voice Separation with an Unknown Number of Multiple Speakers In which, we present a new method for separating a mixed audio sequence, in which multiple voices speak simultaneously. The new method employs gated neural networks that are trained to separate the voices at multiple processing steps, while maintaining the speaker in each output channel fixed. A different model is trained for every number of possible speakers, and the model with the largest number of speakers is employed to select the actual number of speakers in a given sample. Our method greatly outperforms the current state of the art, which, as we show, is not competitive for more than two speakers.
★ 1.3ktext_nn. Text classification models. Used a submodule for other projects.
★ 69PyTorch-VAE. A Collection of Variational Autoencoders (VAE) in PyTorch.
★ 7.7klightly. A python library for self-supervised learning on images.
★ 3.8kgdown.pl. Google Drive direct download of big files
★ 944ResDAVEnet-VQ. Official codes for the paper "Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech"
★ 28End-to-end-E2E-Named-Entity-Recognition-from-English-Speech. Python
★ 32UnsupSeg. Self-Supervised Contrastive Learning for Unsupervised Phoneme Segmentation (INTERSPEECH 2020)
★ 147EfficientNet-PyTorch. A PyTorch implementation of EfficientNet
★ 8.2kentity-recognition-datasets. A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of languages, domains and entity types.
★ 1.6kGAN_Harmonized_with_HMMs. Code:Completely Unsupervised Speech Recognition By A Generative Adversarial Network Harmonized With Iteratively Refined Hidden Markov Models
★ 25neural_manifolds_replicaMFT. Python
★ 81pet. This repository contains the code for "Exploiting Cloze Questions for Few-Shot Text Classification and Natural Language Inference"
★ 1.6k