pystoi. Python implementation of the Short Term Objective Intelligibility measure
359pytorch_stoi. STOI loss function in PyTorch
107pywsj0-mix. wsj0-{2, 3, 4, 5} mix generation scripts, in Python.
79Ranger-Deep-Learning-Optimizer. Ranger - a synergistic optimizer using RAdam (Rectified Adam) and LookAhead in one codebase
12Speech-Separation-Paper. A must-read paper for speech separation based on neural networks
9pb_bss. Collection of EM algorithms for blind source separation of audio signals
6Intelligibility-MetricGAN. Implementation for paper "iMetricGAN: Intelligibility Enhancement for Speech-in-Noise using Generative Adversarial Network-based Metric Learning"
3jn-scripts. Common things to do for Jetson Nano DTK
3audio. simple audio I/O for pytorch
3seewav. Audio waveform visualisation, converts any audio to a nice video
3mpariente.
3denoiser. Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities.
3DNS-Challenge. This repo contains the scripts, models and required files for the Interspeech 2020 Deep Noise Suppression (DNS) Challenge. We are open sourcing clean speech and noise files as well. Participants of this challenge will use the scripts from this repo to create data to train their noise suppressors. They will compare their method with our baseline noise suppressor and report the results.
2Best-websites-a-programmer-should-visit. :link: Some useful websites for programmers.
2q. C++ Library for Audio Digital Signal Processing
2transformers. 🤗Transformers: State-of-the-art Natural Language Processing for Pytorch and TensorFlow 2.0.
2SoundCard. A Pure-Python Real-Time Audio Library
2datasets. 🤗 Fast, efficient, open-access datasets and evaluation metrics in PyTorch, TensorFlow, NumPy and Pandas
1mpariente.github.io. ✨ Build a beautiful and simple website in literally minutes.
1pytorch-lightning. The lightweight PyTorch wrapper for ML researchers. Scale your models. Write less boilerplate
1speechmetrics. A wrapper around speech quality metrics MOSNet, BSSEval, STOI, PESQ, SRMR
1simple-cython-limiter. A simple real-time limiter implemented using Python, Cython, Numpy and PyAudio
1torchsample. High-Level Training, Data Augmentation, and Utilities for Pytorch
1sound-separation. Python
1markdown_readme. Markdown - you can mark up titles, lists, tables, etc., in a much cleaner, readable and accurate way if you do it with HTML.
1librosa. Python library for audio and music analysis
1soxbindings. Python bindings for SoX, aiming to replicate a subset of the command line sox utility.
1jiwer. Evaluate your speech-to-text system with similarity measures such as word error rate (WER)
1buildKernelAndModules. Build the Linux Kernel and Modules on board the NVIDIA Jetson Nano Developer Kit
1xling-SemDiv. Code and data for the EMNLP 2020 paper: "Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to Rank"
1pyannote-audio. Neural building blocks for speaker diarization: speech activity detection, speaker change detection, speaker embedding
1OrdNMF. Python
1torchcrepe. Pytorch implementation of the CREPE pitch tracker
1