This is your work, valued
PhD student working on audio, speech, language, and brain. OHISAMA.
Mamba-TasNet. Jupyter Notebook
★ 116Mamba-ASR. ConMamba for Automatic Speech Recognition
★ 106Style-Talker. An official implementation of Style-Talker for Spoken Dialogue Generation
★ 23AAD-LLM. AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
★ 2ECE391-Kernel. C
★ 1ListenChatRemix. Source code for ListenChatRemix
★ 1streamfm. Real-Time Streamable Generative Speech Restoration with Flow Matching
★ 60unified-audio. An Open-Source Project to Unify Audio Processing and Generation
★ 484moshi-finetune. Python
★ 475glm-4-voice-finetune. Python
★ 14Awesome-Video-Generation-Post-Training. [TMLR] Video Generation Models: A Survey of Post-Training and Alignment | 🔥 A continuously updated collection of papers, datasets, and benchmarks on post-training and alignment for video generation.
★ 173wjk0925.github.io. HTML
★ 1Zonos. Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers.
★ 7.2kWSI. Whisper Speaker Identification (WSI), a cutting-edge model for multilingual speaker identification.
★ 27F5-TTS. Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
★ 15kslamkit. SlamKit is an open source tool kit for efficient training of SpeechLMs. It was used for "Slamming: Training a Speech Language Model on One GPU in a Day"
★ 229audiobox-aesthetics. Unified automatic quality assessment for speech, music, and sound.
★ 749stable-codec. A family of state-of-the-art Transformer-based audio codecs for low-bitrate high-quality audio coding.
★ 438Machine-Learning-From-Scratch. Implementation of popular ML algorithms from scratch
★ 981disentangling_representations. Python
★ 14SoloAudio. SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.
★ 121SemantiCodec-inference. Ultra-low bitrate neural audio codec (0.31~1.40 kbps) with a better semantic in the latent space.
★ 254MuCodec. Python
★ 169Qwen2-Audio. The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 40moshi. Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
★ 11kaudio-diffusion-pytorch. Audio generation using diffusion models, in PyTorch.
★ 2.1kStyleTTS-ZS. StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
★ 188llm-hallucination-survey. Reading list of hallucination in LLMs. Check out our new survey paper: "Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models"
★ 1.1kBayLing-Speech. LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
★ 3.1kMSLDM. Implementation of Multi-Source Music Generation with Latent Diffusion.
★ 29slt2024-speechllm-speaker-understanding.
★ 7xcodec. AAAI 2025: Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
★ 308BigCodec. Official implementation of the paper "BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec"
★ 218audio-ai-hub. The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
★ 949Mamba-TasNet. Jupyter Notebook
★ 116autoregressive-diffusion-pytorch. Implementation of Autoregressive Diffusion in Pytorch
★ 438audio-flamingo. PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
★ 1.2kssamba. [SLT'24] The official implementation of SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
★ 141Qwen2-Audio. The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 2.1kMamba-ASR. ConMamba for Automatic Speech Recognition
★ 106ms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kAudioFlamingo. Implementation of the model "AudioFlamingo" from the paper: "Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities"
★ 39Awesome-Mamba. Awesome list of papers that extend Mamba to various applications.
★ 142mamba. Mamba SSM architecture
★ 19kmamba-minimal. Simple, minimal implementation of the Mamba SSM in one file of PyTorch.
★ 3kICLR-2024-OpenReview-Ratings.
★ 31Qwen-Audio. The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 1.9kMusic-Demixing-with-Band-Split-RNN. An unofficial PyTorch implementation of Music Source Separation with Band-split RNN for MDX-23 ("Label Noise" Track)
★ 201awesome-large-audio-models. Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
★ 735ltu. Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
★ 478UniAudio. The Open Source Code of UniAudio
★ 606SpeechT5. Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
★ 1.5kStyleTTS2. StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
★ 6.3kLASS. This repo hosts the code and model of "Separate What You Describe: Language-Queried Audio Source Separation", Interspeech 2022
★ 146CLAP. Contrastive Language-Audio Pretraining
★ 2.2kclap_curation. Python
★ 6CLAP. Learning audio concepts from natural language supervision
★ 673SpatialCodec. Implementation of SpatialCodec.
★ 71VALL-E-X. An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
★ 7.9kvall-e. An unofficial PyTorch implementation of the audio LM VALL-E
★ 3kvocos. Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
★ 1.1kSpeechGPT. SpeechGPT Series: Speech Large Language Models
★ 1.4kNExT-GPT. Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6ktextlesslib. Library for Textless Spoken Language Processing
★ 559ODSQA. ODSQA: OPEN-DOMAIN SPOKEN QUESTION ANSWERING DATASET
★ 64Awesome-LLM. Awesome-LLM: a curated list of Large Language Model
★ 27kVGGSound. VGGSound: A Large-scale Audio-Visual Dataset
★ 359ml-spatial-librispeech. A large synthetic dataset of spatial audio with multiple labels
★ 127llama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
★ 19kunilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kllama. Inference code for Llama models
★ 60kvisqol. Perceptual Quality Estimator for speech and audio
★ 913audiolm-pytorch. Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch
★ 2.6kTwo-Stage-Polyphonic-Sound-Event-Detection-and-Localization. A two-stage polyphonic sound event detection and localization method for both SED and DOA.
★ 126encodec-pytorch. unofficial implementation of the High Fidelity Neural Audio Compression
★ 176TAC. transform-average-concatenate (TAC) method for end-to-end microphone permutation and number invariant ad-hoc beamforming.
★ 311gpuRIR. Python library for Room Impulse Response (RIR) simulation with GPU acceleration
★ 607DCASE2022-data-generator. Data generator for creating synthetic audio mixtures suitable for DCASE Challenge 2022 Task 3
★ 47uclser20. Code for the paper "Unsupervised Contrastive Learning of Sound Event Representations", ICASSP 2021.
★ 93seld-dcase2023. Baseline method for sound event localization task of DCASE 2023 challenge
★ 71vocoder-benchmark. A repository for benchmarking neural vocoders by their quality and speed.
★ 213consistency_models. Official repo for consistency models.
★ 6.5kstable-diffusion. A latent text-to-image diffusion model
★ 73khifi-gan. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
★ 2.4kSoundStream. This repository is an implementation of this article: https://arxiv.org/pdf/2107.03312.pdf
★ 431NeuralSpeech. Python
★ 1.5kaudio-diffusion. Apply diffusion models using the new Hugging Face diffusers package to synthesize music instead of images.
★ 792diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kAwesome-Diffusion-Models. A collection of resources and papers on Diffusion Models
★ 12kFullySpikingVAE. Official implementation of Fully Spiking Variational Autoencoder [AAAI2022]
★ 75espnet. End-to-End Speech Processing Toolkit
★ 9.9ks3prl. Self-Supervised Speech Pre-training and Representation Learning Toolkit
★ 2.6kavalanche. Avalanche: an End-to-End Library for Continual Learning based on PyTorch.
★ 2.1kKnowledge-Distillation-Zoo. Pytorch implementation of various Knowledge Distillation (KD) methods.
★ 1.8kSPQ. Self-supervised Product Quantization for Deep Unsupervised Image Retrieval - ICCV2021
★ 87pyloudnorm. Flexible audio loudness meter in Python with implementation of ITU-R BS.1770-4 loudness algorithm
★ 776Conv-TasNet. Python
★ 337Discriminator-Constrained-Optimal-Transport-Network. Python
★ 30sound-separation. Python
★ 720cssl_sound. Python
★ 14barlowtwins. PyTorch implementation of Barlow Twins.
★ 1kpytorch-metric-learning. The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in PyTorch.
★ 6.3kDANN. pytorch implementation of Domain-Adversarial Training of Neural Networks
★ 951awesome-domain-adaptation. A collection of AWESOME things about domain adaptation
★ 5.5kspeechbrain. A PyTorch-based Speech Toolkit
★ 12kBBAVectors-Oriented-Object-Detection. [WACV2021] Oriented Object Detection in Aerial Images with Box Boundary-Aware Vectors
★ 474pytorch-vqvae. Vector Quantized VAEs - PyTorch Implementation
★ 955vector-quantize-pytorch. Vector (and Scalar) Quantization, in Pytorch
★ 4klinux. Linux kernel source tree
★ 241kdenoiser. Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities.
★ 1.9kpytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
★ 102kECE391-Kernel. C
★ 2