This is your work, valued
Super user, wisdom giver, and bacon & eggs master. Working on Natural General Stupidity.
Multilingual_Text_to_Speech. An implementation of Tacotron 2 that supports multilingual experiments with parameter-sharing, code-switching, and voice cloning.
★ 844Star_Tracker. Arduino DIY telescope GoTo for arbitrary mounts.
★ 66MultiWOZ_Evaluation. Unified MultiWOZ evaluation scripts for the context-to-response task.
★ 59Blizzard2013_Segmentation. Transcripts and segmentation for the Blizzard 2013 audiobooks also known as the Lessac or Blizzard 2013 dataset.
★ 45Aargh. Python
★ 12WaveRNN. WaveRNN Vocoder + TTS
★ 11Unity_Tower_Defence. Classical top-down tower defence made in Unity 3D game engine.
★ 9Face_Cleaner. Automated trimming and cleaning of 3D facial scans
★ 4UE4_Endless_Racer. Endless racer (runner) created using Blueprints Visual Scripting system of the Unreal Engine 4.
★ 3Toom_Rendering_Engine. Partial remake of the original Doom 1. Written as an assignment during a programming course. Uses doom-like rendering.
★ 3Sequicity_Knowledge_Base. Implementation of knowledge base for the sequicity model.
★ 1Pascal_Star_Fighter. A simple Star Fighter game which I created as a final assignment in the introductory course of programming during the first semester at the uni.
★ 1tomiinek.github.io. HTML
★ 1Phaser3_Space_Shooter. A simple 2D shooter exploiting features of the Phaser 3 framework.
★ 1MOSS-TTS. MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
★ 3.9kadversarial-robustness-toolbox. Adversarial Robustness Toolbox (ART) - Python Library for Machine Learning Security - Evasion, Poisoning, Extraction, Inference - Red and Blue Teams
★ 6.1kOmniVoice. High-Quality Voice Cloning TTS for 600+ Languages
★ 8.7kPeriodWave. The official Implementation of PeriodWave and PeriodWave-Turbo
★ 226csm. A Conversational Speech Generation Model
★ 15kFlashLabs-Chroma. Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.
★ 550Qwen3-TTS. Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
★ 13kRespiro-en. Official implementation of paper: Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
★ 44LinaCodec. A highly compressive and high-quality neural audio codec for speech models.
★ 269cro-dl. 🐍 CRo-DL (Český Rozhlas Downloader) - MůjRozhlas.cz 📻 offline
★ 2tts. Inworld TTS
★ 736VoxCPM. VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
★ 35kX-Codec-2.0. Codec for paper: LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
★ 362chatterbox. SoTA open-source TTS
★ 26kSidon. Training code and dataset cleasing with Sidon
★ 172adversarial-attacks-pytorch. PyTorch implementation of adversarial attacks [torchattacks]
★ 2.2kEmergentTTS-Eval-public. [NeurIPS' 25] Benchmark for evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges.
★ 226S3Tokenizer. Reverse Engineering of Supervised Semantic Speech Tokenizer (S3Tokenizer) proposed in CosyVoice
★ 521delayed-streams-modeling. Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.
★ 3kBlaGPT. Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration.
★ 113tidy-tunes. Tidy Tunes is an easy-to-use pipeline for mining high-quality audio data for speech generation models. To do so, it chains multiple open source models while minimizing dependencies.
★ 23pytorch-nips2017-attack-example. A PyTorch baseline attack example for the NIPS 2017 adversarial competition
★ 87ten-vad. Voice Activity Detector (VAD) : low-latency, high-performance and lightweight
★ 2.2kparsonaut. Python
★ 4Kimi-Audio. Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
★ 4.7kdia. A TTS model capable of generating ultra-realistic dialogue in one pass.
★ 19kOrpheus-TTS. Towards Human-Sounding Speech
★ 6.3kZonos. Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers.
★ 7.2kSpeech-Articulatory-Coding. Jupyter Notebook
★ 67audiobox-aesthetics. Unified automatic quality assessment for speech, music, and sound.
★ 749open-r1. Fully open reproduction of DeepSeek-R1
★ 26kMatcha-TTS. [ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching
★ 1.3kcuml. cuML - RAPIDS Machine Learning Library
★ 5.2kawesome-maps. There is more than google: A collection of great online maps 🌍🗺🌎
★ 500biblically-accurate-sampler. llm sampler that only allows words that are in the bible
★ 43stable-ts. Transcription, forced alignment, and audio indexing with OpenAI's Whisper
★ 2.3kaudiotools. Object-oriented handling of audio data, with GPU-powered augmentations, and more.
★ 350Fast-GeCo. Source code and demo for INTERSPEECH 2024 paper: Noise-robust Speech Separation with Fast Generative Correction
★ 50gpt-bert. Official implementation of "GPT or BERT: why not both?"
★ 64Transformer-Explainability. [CVPR 2021] Official PyTorch implementation for Transformer Interpretability Beyond Attention Visualization, a novel method to visualize classifications by Transformer based networks.
★ 2kseed-vc. zero-shot voice conversion & singing voice conversion, with real-time support
★ 3.9kF5-TTS. Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
★ 15kAdafruit_TSL2561. Unified sensor driver for Adafruit's TSL2561 breakouts
★ 89wavefit-pytorch. PyTorch implementation of WaveFit [2022, Google] which is one of SOTA lightweight/fast speech vocoders.
★ 70moshi. Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
★ 11kdemucs. Code for the paper Hybrid Spectrogram and Waveform Source Separation
★ 3kWavTokenizer. [ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
★ 1.3kutmos. A toolkit to calculate speech audio quality. Not affiliated with the original authors
★ 74SpeechTokenizer. This is the code for the SpeechTokenizer presented in the SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models. Samples are presented on
★ 658FreeV. [InterSpeech 24] FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
★ 98torchdiffeq. Differentiable ODE solvers with full GPU support and O(1)-memory backpropagation.
★ 6.5kLibriTTS-P. LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
★ 162friendly-stable-audio-tools. Refactored / updated version of `stable-audio-tools` which is an open-source code for audio/music generative models originally by Stability AI.
★ 218astral. ☀️ Go calculations for the position of the sun and moon.
★ 35awesome-discrete-diffusion-models. A curated list for awesome discrete diffusion models resources.
★ 572tts-arabic-pytorch. 🎙️ Arabic TTS models (Tacotron2, FastPitch)
★ 145arabic_vocalizer. Arabic deep-learning based diacritization models (Shakkala, Shakkelha) in the ONNX format.
★ 12snac. Multi-Scale Neural Audio Codec (SNAC) compresses audio into discrete codes at a low bitrate
★ 774dataspeech. Python
★ 400parler-tts. Inference and training library for high-quality TTS models.
★ 5.6kdasp-pytorch. Differentiable audio signal processors in PyTorch
★ 298torch-pitch-shift. Pitch-shift audio clips quickly with PyTorch (CUDA supported)! Additional utilities for searching efficient transformations are included.
★ 139metavoice-src. Foundational model for human-like, expressive TTS
★ 4.2kDeepFilterNet. Noise supression using deep filtering
★ 4.5kdescript-audio-codec. State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.
★ 1.8kgpt-fast. Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.
★ 6.2kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24kpyctcdecode. A fast and lightweight python-based CTC beam search decoder for speech recognition.
★ 469ControlNet. Let us control diffusion models!
★ 34kmamba. Mamba SSM architecture
★ 19kttts. Train the next generation of TTS systems.
★ 169vocos. Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
★ 1.1kFreeVC. FreeVC: Towards High-Quality Text-Free One-Shot Voice Conversion
★ 714national-anthems-clustering. Jupyter Notebook
★ 90ProphetNet. A research project for natural language generation, containing the official implementations by MSRA NLC team.
★ 746diffusion. Denoising Diffusion Probabilistic Models
★ 5.3kvocode-core. 🤖 Build voice-based LLM agents. Modular + open source.
★ 3.8klaughter-detection. Python
★ 291BigVGAN. BigVGAN with Neural Source-Filter
★ 58StarGANv2-VC. StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
★ 522AutoVocoder. Autovocoder: Fast Waveform Generation from a Learned Speech Representation using Differentiable Digital Signal Processing
★ 71Knowledge-Distillation-Toolkit. :no_entry: [DEPRECATED] A knowledge distillation toolkit based on PyTorch and PyTorch Lightning.
★ 138unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kpanns_inference. Python
★ 266fairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
★ 32kCLAP. Contrastive Language-Audio Pretraining
★ 2.2ktortoise-tts-fast. Fast TorToiSe inference (5x or your money back!)
★ 826so-vits-svc-fork. so-vits-svc fork with realtime support, improved interface and more features.
★ 9.3kencodec. State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
★ 4kconsistency_models. Official repo for consistency models.
★ 6.5kPrompt-Engineering-Guide. 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
★ 77kLarge-Audio-Models. Keep track of big models in audio domain, including speech, singing, music etc.
★ 516univnet. Unofficial PyTorch Implementation of UnivNet Vocoder (https://arxiv.org/abs/2106.07889)
★ 286torch-nansypp. NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis
★ 152SimpleParsing. Simple, Elegant, Typed Argument Parsing with argparse
★ 543libebur128. A library implementing the EBU R128 loudness standard.
★ 490simple-asgan. Training code and trained checkpoints for ASGAN.
★ 62prompt-to-prompt. Jupyter Notebook
★ 3.5kHarmoF0. Python
★ 108wav2lip-hq. Extension of Wav2Lip repository for processing high-quality videos.
★ 547whisper. Robust Speech Recognition via Large-Scale Weak Supervision
★ 106kdenoiser. Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities.
★ 1.9kAwesome-Singing-Voice-Synthesis-and-Singing-Voice-Conversion. A paper and project list about the cutting edge Speech Synthesis, Text-to-Speech (TTS), Singing Voice Synthesis (SVS), Voice Conversion (VC), Singing Voice Conversion (SVC), and related interesting works (such as Music Synthesis, Automatic Music Transcription, Automatic MOS Prediction, SSL-based ASR...etc).
★ 488fast_pytorch_kmeans. This is a pytorch implementation of k-means clustering algorithm
★ 346TorchPQ. Approximate nearest neighbor search with product quantization on GPU in pytorch and cuda
★ 237pyxmeans. Quick implementation of xmeans in python and C
★ 87pedalboard. 🎛 🔊 A Python library for audio.
★ 6.2kDailyTalk. Official repository of DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech, ICASSP 2023
★ 259ai-audio-startups. Community list of startups working with AI in audio and music technology
★ 1.8kcuda-samples. Samples for CUDA Developers which demonstrates features in CUDA Toolkit
★ 9.4kkeops. KErnel OPerationS, on CPUs and GPUs, with autodiff and without memory overflows
★ 1.2kcylimiter. A C++/Cython audio limiter for Python.
★ 25diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kPython-Wrapper-for-World-Vocoder. A Python wrapper for the high-quality vocoder "World"
★ 790ParaLip. Parallel and High-Fidelity Text-to-Lip Generation; AAAI 2022 ; Official code
★ 109textlesslib. Library for Textless Spoken Language Processing
★ 559nvidia-htop. A tool for enriching the output of nvidia-smi.
★ 580NLP-progress. Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the most common NLP tasks.
★ 23klyra. A Very Low-Bitrate Codec for Speech Compression
★ 4kTransformerTTS. 🤖💬 Transformer TTS: Implementation of a non-autoregressive Transformer based neural network for text to speech.
★ 1.2kvector-quantize-pytorch. Vector (and Scalar) Quantization, in Pytorch
★ 4kqsubmit. Wrapper for batch engine submission commands (remake of Obo's Perl qsubmit)
★ 5torch-imle. Implicit MLE: Backpropagating Through Discrete Exponential Family Distributions
★ 260pytorch-normalizing-flows. Normalizing flows in PyTorch. Current intended use is education not production.
★ 917multilexnorm2021. MultiLexNorm 2021 competition system from ÚFAL
★ 16glow. Code for reproducing results in "Glow: Generative Flow with Invertible 1x1 Convolutions"
★ 3.2kCommonsense-Dialogues. A crowdsourced dataset of dialogues grounded in social contexts involving utilization of commonsense.
★ 80gecko. Gecko - A Tool for Effective Annotation of Human Conversations
★ 306multi-speaker-tacotron. VCTK multi-speaker tacotron for ICASSP 2020
★ 266soloist. Python
★ 77crepe. CREPE: A Convolutional REpresentation for Pitch Estimation -- pre-trained model (ICASSP 2018)
★ 1.4kuda. Unsupervised Data Augmentation (UDA)
★ 2.2kDiverse-Reference-Augmentation. Code and Data for our Findings of ACL 2021 paper titled 'Improving Automated Evaluation of Open Domain Dialog via Diverse Reference Augmentation. Varun Gangal *, Harsh Jhamtani *, Eduard Hovy, Taylor Berg-Kirkpatrick'
★ 6GEM-metrics. Automatic metrics for GEM tasks
★ 69Soft-DTW-Loss. PyTorch implementation of Soft-DTW: a Differentiable Loss Function for Time-Series in CUDA
★ 150vits. VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
★ 7.9kufal-template-alt. Customized UFAL poster template
★ 2lhotse. Tools for handling multimodal data in machine learning projects.
★ 1.1ksurvey. A Survey on Neural Speech Synthesis https://arxiv.org/pdf/2106.15561.pdf
★ 371adapters. A Unified Library for Parameter-Efficient and Modular Transfer Learning
★ 2.8kx-transformers. A concise but complete full-attention transformer with a set of promising experimental features from various papers
★ 5.9kbertviz. BertViz: Visualize Attention in Transformer Models
★ 8.1knopdb. NoPdb: Non-interactive Python Debugger
★ 84DialogWAE. Source Code for DialogWAE: Multimodal Response Generation with Conditional Wasserstein Autoencoder (https://arxiv.org/abs/1805.12352)
★ 126posterior-collapse-list. A curated list of techniques to avoid posterior collapse
★ 92DialogRPT. EMNLP 2020: "Dialogue Response Ranking Training with Large-Scale Human Feedback Data"
★ 345entmax. The entmax mapping and its loss, a family of sparse softmax alternatives.
★ 474sentence-transformers. State-of-the-Art Embeddings, Retrieval, and Reranking
★ 19kfed. Code for SIGdial 2020 paper: Unsupervised Evaluation of Interactive Dialog with DialoGPT (https://arxiv.org/abs/2006.12719)
★ 28usr. Code for ACL 2020 paper: USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation (https://arxiv.org/pdf/2005.00456)
★ 50self_dialogue_corpus. The Self-dialogue Corpus - a collection of self-dialogues across music, movies and sports
★ 107e2e_dialog_challenge. End-To-End Task-Completion Dialogue Challenge
★ 194dstc8-schema-guided-dialogue. The Schema-Guided Dialogue Dataset
★ 608airdialogue. Python
★ 47chat_corpus. chat corpus collection from various open sources
★ 137Topical-Chat. A dataset containing human-human knowledge-grounded open-domain conversations.
★ 673Taskmaster. Please see the readme file as well as our 2019 EMNLP paper linked here -->
★ 222pytorch-lightning. Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
★ 31ksam. SAM: Sharpness-Aware Minimization (PyTorch)
★ 2kTC-Bot. User Simulation for Task-Completion Dialogues
★ 802simmc. With the aim of building next generation virtual assistants that can handle multimodal inputs and perform multimodal actions, we introduce two new datasets (both in the virtual shopping domain), the annotation schema, the core technical tasks, and the baseline models. The code for the baselines and the datasets will be opensourced.
★ 133einops. Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)
★ 9.6klambda-networks. Implementation of LambdaNetworks, a new approach to image recognition that reaches SOTA with less compute
★ 1.5klambda-bert. A 🤗-style implementation of BERT using lambda layers instead of self-attention
★ 69faiss. A library for efficient similarity search and clustering of dense vectors.
★ 41kunlikelihood_training. Neural Text Generation with Unlikelihood Training
★ 311ToD-BERT. Pre-Trained Models for ToD-BERT
★ 295Target-Guided-Conversation. "Target-Guided Open-Domain Conversation" in ACL 2019
★ 147ParlAI. A framework for training and evaluating AI models on a variety of openly available dialogue datasets.
★ 11kdialog-rl. Python
★ 3perin. PERIN is Permutation-Invariant Semantic Parser developed for MRP 2020
★ 45howdy. 🛡️ Windows Hello™ style facial authentication for Linux
★ 7.7kada-hessian. Easy-to-use AdaHessian optimizer (PyTorch)
★ 79SC-GPT. Few-shot Natural Language Generation for Task-Oriented Dialog
★ 190ConvLab-2. ConvLab-2: An Open-Source Toolkit for Building, Evaluating, and Diagnosing Dialogue Systems
★ 466NeuralDialogPapers. Summary of deep learning models for dialog systems (Tiancheng Zhao LTI, CMU)
★ 643BRC. A repository containing the code for the Bistable Recurrent Cell
★ 47Cross-Lingual-Voice-Cloning. Tacotron 2 - PyTorch implementation with faster-than-realtime inference modified to enable cross lingual voice cloning.
★ 359LeetCode. :pencil2: LeetCode solutions in C++ 11 and Python3
★ 3.5kdetr. End-to-End Object Detection with Transformers
★ 15kespnet. End-to-End Speech Processing Toolkit
★ 9.9kspeedyspeech. Python
★ 262WaveRNN. WaveRNN Vocoder + TTS
★ 2.2kmeta-tasnet. A PyTorch implementation of Meta-TasNet from "Meta-learning Extractors for Music Source Separation
★ 138lazy-adam. :angel: Lazy Adam optimizer (pytorch)
★ 5datasets-CMU_Wilderness. CMU Wilderness Multilingual Speech Dataset
★ 292AdaIN-style. Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization
★ 1.6kcss10. CSS10: A Collection of Single Speaker Speech Datasets for 10 Languages
★ 490tacotron. A TensorFlow implementation of Google's Tacotron speech synthesis with pre-trained model (unofficial)
★ 3kTTS. :robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
★ 10kReal-Time-Voice-Cloning. Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 60ktgen. Statistical NLG for spoken dialogue systems
★ 207TinyEKF. Lightweight C/C++ Extended Kalman Filter with Python for prototyping
★ 1.2ki2cdevlib. I2C device library collection for AVR/Arduino or other C++-based MCUs
★ 4.3kKalmanFilter. This is a Kalman filter used to calculate the angle, rate and bias from from the input of an accelerometer/magnetometer and a gyroscope.
★ 1.9kneat-openai-gym. NEAT for Reinforcement Learning on the OpenAI Gym
★ 26CEM-RL. Combining Evolutionary Algorithms and deep RL in various ways
★ 108ArduinoMotionSensorExample. MPU6050/MPU6500/MPU9150/MPU9250 over I2c for Arduino
★ 120codingame-cpp-merge. Automatic merger for C/C++ in the CodinGame IDE
★ 19deep-voice-conversion. Deep neural networks for voice conversion (voice style transfer) in Tensorflow
★ 3.9k