This is your work, valued
Interested in sound generation. Let's make AI more powerful.
FastSpeech. The Implementation of FastSpeech based on pytorch.
★ 885FastVocoder. Include Basis-MelGAN, MelGAN, HifiGAN and Multiband-HifiGAN, maybe NHV in the future.
★ 157Transformer-TTS. TTS model based on Transformer.
★ 57FastSpeech2. The Implementation of FastSpeech2 Based on Pytorch.
★ 52ConvTasNet4BasisMelGAN. This repo contains conv-tasnet for basis-melgan. If you want to get code of basis-melgan, please refer to FastVocoder.
★ 21CLONE.
★ 20Tacotron2-Pytorch. follow NVIDIA, simplify it and support data parallel.
★ 13Lifelong-Learning-Tacotron2. MultiSpeaker Tacotron2 using LifeLong Learning.
★ 13tacotron2.xcmyz. new version of tacotron2 (old version: https://github.com/xcmyz/Tacotron2-Pytorch)
★ 8Hackathon-EnglishLearning. Voice Scoring System.
★ 8LM-Tacotron2. Tacotron2 Combine with Language Model (BERT).
★ 7SpeakerVerification. Speaker Verification (GE2E Loss)
★ 7Gobang-AI. A C++ Implementation of Gobang AI.
★ 6Speech-Resources. 语音方向实验室/公司/资源/实习等,欢迎推荐或自荐(排名不分先后)
★ 3Forced-Alignment. using montreal-forced-aligner.
★ 2bert-race. BERT/ALBERT based model for RACE dataset, support multi-worker, multi-GPU, FP16 and bind CPU.
★ 2AVX-programming. CPU acceleration using AVX (Advanced Vector Extensions)
★ 1xcmyz.
★ 1ExpressionTransformation. prefix expression, infix expression, postfix expression.
★ 1diffwave. DiffWave is a fast, high-quality neural vocoder and waveform synthesizer.
★ 1Large-Audio-Models. Keep track of big models in audio domain, including speech, singing, music etc.
★ 1Calculator. A Calculator implemented in Python.
★ 1VAE-Tacotron. A Pytorch Implementation of Tacotron Combined with VAE
★ 1FaceDetection. Python
★ 1Polynomial-Calculator. 基于Python实现的带有图形界面的多项式计算器
★ 1x-transformers. A concise but complete full-attention transformer with a set of promising experimental features from various papers
★ 5.9kseed-vc. zero-shot voice conversion & singing voice conversion, with real-time support
★ 3.9kLLaSA_training. LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
★ 661DeepSeek-R1.
★ 92kmusic_llm. Python
★ 56mini_llm. Python
★ 29VITA. ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
★ 2.5kmoshi. Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
★ 11kmini-omni. open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
★ 3.6kWavTokenizer. [ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
★ 1.3kdescript-audio-codec. State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.
★ 1.8kllama-models. Utilities intended for use with Llama models.
★ 7.7kllama3. The official Meta Llama 3 GitHub site
★ 29kfish-speech. SOTA Open Source TTS
★ 32kvideollm-online. VideoLLM-online: Online Video Large Language Model for Streaming Video (CVPR 2024)
★ 679chameleon. Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
★ 2.1kOpen-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kseed-tts-eval. Python
★ 1.6kbytedancespeech.github.io. HTML
★ 37U-ViT. A PyTorch implementation of the paper "All are Worth Words: A ViT Backbone for Diffusion Models".
★ 1.1kSiT. Official PyTorch Implementation of "SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers"
★ 1.2kLumina-T2X. Lumina-T2X is a unified framework for Text to Any Modality Generation
★ 2.2kHDTF. the dataset and code for "Flow-guided One-shot Talking Face Generation with a High-resolution Audio-visual Dataset"
★ 429Real-ESRGAN. Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
★ 36kOpen-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kmagvit2-pytorch. Implementation of MagViT2 Tokenizer in Pytorch
★ 668EMO. Emote Portrait Alive: Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
★ 7.6kDiT. Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
★ 8.7kdiffae. Official implementation of Diffusion Autoencoders
★ 968GPT-SoVITS. 1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 60kpytorch-vq-vae. PyTorch implementation of VQ-VAE by Aäron van den Oord et al.
★ 604bigvsan. Pytorch implementation of BigVSAN
★ 203ctc-segmentation. Segment an audio file and obtain utterance alignments. (Python package)
★ 348audiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24kflash-attention. Fast and memory-efficient exact attention
★ 25kWavJourney. WavJourney: Compositional Audio Creation with LLMs
★ 544vall-e. PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
★ 2.2kaudio-ai-timeline. A timeline of the latest AI models for audio generation, starting in 2023!
★ 1.9kMeta-voicebox. Implementation of Meta-Voicebox : The first generative AI model for speech to generalize across tasks with state-of-the-art performance.
★ 595Stepwise_Monotonic_Multihead_Attention. PyTorch Implementation of Stepwise Monotonic Multihead Attention similar to Enhancing Monotonicity for Robust Autoregressive Transformer TTS
★ 39Macaw-LLM. Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration
★ 1.6kGenAI_LLM_timeline. ChatGPT, GenerativeAI and LLMs Timeline
★ 953AcademiCodec. AcademiCodec: An Open Source Audio Codec Model for Academic Research
★ 674SpeechGPT. SpeechGPT Series: Speech Large Language Models
★ 1.4kWhisperSpeech. An Open Source text-to-speech system built by inverting Whisper.
★ 4.6kpytorch-lightning. Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
★ 31kLLMSurvey. The official GitHub page for the survey paper "A Survey of Large Language Models".
★ 12kvocalsound. Dataset and baseline code for the VocalSound dataset (ICASSP2022).
★ 165StableLM. StableLM: Stability AI Language Models
★ 16kMOSS. An open-source tool-augmented conversational language model from Fudan University
★ 12kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kMelNet. Implementation of "MelNet: A Generative Model for Audio in the Frequency Domain"
★ 209bark. 🔊 Text-Prompted Generative Audio Model
★ 39kgpt-2. Code for the paper "Language Models are Unsupervised Multitask Learners"
★ 25kmusiclm-pytorch. Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch
★ 3.3kAwesome-LLM. Awesome-LLM: a curated list of Large Language Model
★ 27kLarge-Audio-Models. Keep track of big models in audio domain, including speech, singing, music etc.
★ 516large_language_model_training_playbook. An open collection of implementation tips, tricks and resources for training large language models
★ 503audiolm-pytorch. Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch
★ 2.6klibri-light. dataset for lightly supervised training using the librivox audio book recordings. https://librivox.org/.
★ 528v-diffusion-pytorch. v objective diffusion inference code for PyTorch.
★ 719disco-diffusion. Jupyter Notebook
★ 7.4kk-diffusion. Karras et al. (2022) diffusion models for PyTorch
★ 2.6kpaper-reading. 深度学习经典、新论文逐段精读
★ 34kAudioLDM. AudioLDM: Generate speech, sound effects, music and beyond, with text.
★ 2.9kChatYuan. ChatYuan: Large Language Model for Dialogue in Chinese and English
★ 1.9kaudio-diffusion. Apply diffusion models using the new Hugging Face diffusers package to synthesize music instead of images.
★ 792vector-quantize-pytorch. Vector (and Scalar) Quantization, in Pytorch
★ 4kEATS. A pytroch implementation of the EETS: End-to-End Adversarial Text-to-Speech
★ 127Speech-Backbones. This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
★ 604Text-to-sound-Synthesis. The source code of our paper "Diffsound: discrete diffusion model for text-to-sound generation"
★ 366MB-iSTFT-VITS. Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
★ 469improved-diffusion. Release for Improved Denoising Diffusion Probabilistic Models
★ 3.8kscore_sde_pytorch. PyTorch implementation for Score-Based Generative Modeling through Stochastic Differential Equations (ICLR 2021, Oral)
★ 2.1kAnalytic-DPM. Code for the paper Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models (ICLR 2022 Outstanding Paper Award)
★ 173iclr2023_stats. ICLR2023 statistics
★ 60Poisson_flow. Code for NeurIPS 2022 Paper, "Poisson Flow Generative Models" (PFGM)
★ 874cycle-diffusion. [ICCV 2023] A latent space for stochastic diffusion models
★ 658NANSY. Python
★ 171GON. Gradient Origin Networks - a new type of generative model that is able to quickly learn a latent representation without an encoder
★ 160encodec. State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
★ 4kMSMC-TTS. Official Implement of Multi-Stage Multi-Codebook (MSMC) TTS
★ 168vdm. Jupyter Notebook
★ 331SpeechT5. Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
★ 1.5kAdaVocoder. Adaptive Vocoder for Custom Voice
★ 61ProDiff. PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
★ 432interspeech2022.
★ 162ddim. Denoising Diffusion Implicit Models
★ 1.8kibot. iBOT :robot:: Image BERT Pre-Training with Online Tokenizer (ICLR 2022)
★ 778FastDeploy. High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7kucxpp. C++
★ 21cs228-notes. Course notes for CS228: Probabilistic Graphical Models.
★ 2kaudio-diffusion-pytorch. Audio generation using diffusion models, in PyTorch.
★ 2.1kstable-diffusion. A latent text-to-image diffusion model
★ 73kawesome-drones-zh. 无人机资源汇总
★ 608VQ-Diffusion. Official implementation of VQ-Diffusion
★ 982moco. PyTorch implementation of MoCo: https://arxiv.org/abs/1911.05722
★ 5.1kblur-diffusion. Official PyTorch implementation of the paper Progressive Deblurring of Diffusion Models for Coarse-to-Fine Image Synthesis.
★ 157LibtorchTutorials. This is a code repository for pytorch c++ (or libtorch) tutorial.
★ 839openit. 致力于打造免费无感的翻墙环境
★ 2.2kdenoising-diffusion-pytorch. Implementation of Denoising Diffusion Probabilistic Model in Pytorch
★ 11kSYsU-lang. A mini, simple and modular compiler for SYsU/SysY(tiny C). Based on Clang/LLVM/ANTLR4/Bison/Flex.
★ 219VQ-Diffusion. Python
★ 487AudioMAE. This repo hosts the code and models of "Masked Autoencoders that Listen".
★ 673diffusion_models. Minimal standalone example of diffusion model
★ 163SummaryOfLoanSuspension. 全国各省市停贷通知汇总
★ 20kAvocodo-pytorch. Avocodo: Generative Adversarial Network for Artifact-free Vocoder
★ 122WaveGrad. Implementation of WaveGrad high-fidelity vocoder from Google Brain in PyTorch.
★ 409CLONE.
★ 20Comprehensive-E2E-TTS. A Non-Autoregressive End-to-End Text-to-Speech (text-to-wav), supporting a family of SOTA unsupervised duration modelings. This project grows with the research community, aiming to achieve the ultimate E2E-TTS
★ 147denoising-diffusion-gan. Tackling the Generative Learning Trilemma with Denoising Diffusion GANs https://arxiv.org/abs/2112.07804
★ 759radtts. Provides training, inference and voice conversion recipes for RADTTS and RADTTS++: Flow-based TTS models with Robust Alignment Learning, Diverse Synthesis, and Generative Modeling and Fine-Grained Control over of Low Dimensional (F0 and Energy) Speech Attributes.
★ 291tortoise-tts. A multi-voice TTS system trained with an emphasis on quality
★ 15kwetts. Production First and Production Ready End-to-End Text-to-Speech Toolkit
★ 416BigVGAN. Official PyTorch implementation of BigVGAN (ICLR 2023)
★ 1.2kResemblyzer. A python package to analyze and compare voices with deep learning
★ 3.3kremote-jobs-in-china. 支持远程办公的中国公司
★ 2.9kiSTFTNet-pytorch. iSTFTNet : Fast and Lightweight Mel-spectrogram Vocoder Incorporating Inverse Short-time Fourier Transform
★ 279icassp2022.
★ 64ReduNet. ReduNet
★ 541metaseq. Repo for external large-scale work
★ 6.6kconformer. [Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
★ 1.1kClariNet. A Pytorch Implementation of ClariNet
★ 293run. 润学全球官方指定GITHUB,整理润学宗旨、纲领、理论和各类润之实例;解决为什么润,润去哪里,怎么润三大问题; 并成为新中国人的核心宗教,核心信念。
★ 32kRecorder. html5 js 录音 mp3 wav ogg webm amr g711a g711u 格式,支持pc和Android、iOS部分Web浏览器、Hybrid App(提供Android iOS App源码)、微信,提供ASR语音识别转文字 H5版语音通话聊天示例 DTMF编码解码
★ 5.6knix-tts. 🐤 Nix-TTS: Lightweight and End-to-end Text-to-Speech via Module-wise Distillation
★ 265ssr_eval. Evaluation and Benchmarking of Speech Super-resolution Methods
★ 157Learn2Sing2.0. Diffusion and Mutual Information-Based Target Speaker SVS by Learning from Singing Teacher
★ 182bddm. BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis
★ 238HelloSilicon. An introduction to ARM64 assembly on Apple Silicon Macs
★ 5khum2song. Hum2Song: Multi-track Polyphonic Music Generation from Voice Melody Transcription with Neural Networks
★ 160NeuralSpeech. Python
★ 1.5kbook-text-to-speech. A book about Text-to-Speech (TTS) in Chinese.
★ 612wespeaker. Research and Production Oriented Speaker Verification, Recognition and Diarization Toolkit
★ 1.4kIMS-Toucan. Controllable and fast Text-to-Speech for over 7000 languages!
★ 2.2kNeuralSVB. Learning the Beauty in Songs: Neural Singing Voice Beautifier; ACL 2022 (Main conference); Official code
★ 461academicpages.github.io. Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.
★ 17kDiffGAN-TTS. PyTorch Implementation of DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
★ 349ailab. Python
★ 5.9ktextlesslib. Library for Textless Spoken Language Processing
★ 559NATSpeech. A Non-Autoregressive Text-to-Speech (NAR-TTS) framework, including official PyTorch implementation of PortaSpeech (NeurIPS 2021) and DiffSpeech (AAAI 2022)
★ 1kSpeech-Resources. 语音方向实验室/公司/资源/实习等,欢迎推荐或自荐
★ 609Meta-TTS. Official repository of https://doi.org/10.1109/TASLP.2022.3167258. More up-to-date code is in "refactor" branch.
★ 192pytorch-revgrad. A minimal pytorch package implementing a gradient reversal layer.
★ 158awesome-normalizing-flows. Awesome resources on normalizing flows.
★ 1.6kav_hubert. A self-supervised learning framework for audio-visual speech
★ 996VAENAR-TTS. PyTorch Implementation of VAENAR-TTS: Variational Auto-Encoder based Non-AutoRegressive Text-to-Speech Synthesis.
★ 74flowseq. Generative Flow based Sequence-to-Sequence Toolkit written in Python.
★ 245pytorch-softdtw-cuda. Fast CUDA implementation of (differentiable) soft dynamic time warping for PyTorch
★ 734notes. Course notes
★ 752sandbox. Play time!
★ 200nice_pytorch. Flow model - NICE
★ 7voicefixer_main. General Speech Restoration
★ 287pytorch-flows. PyTorch implementations of algorithms for density estimation
★ 589efficient_tts. Pytorch implementation of "Efficienttts: an efficient and high-quality text-to-speech architecture"
★ 116cvss. CVSS: A Massively Multilingual Speech-to-Speech Translation Corpus
★ 220bert4keras. keras implement of transformers for humans
★ 5.4kspeech-synthesis-paper. List of speech synthesis papers.
★ 1.1kflow. Keras implement of flow-based models
★ 227glow-tts. A Generative Flow for Text-to-Speech via Monotonic Alignment Search
★ 712Awesome-Diffusion-Models. A collection of resources and papers on Diffusion Models
★ 12kCatch-A-Waveform. Official pytorch implementation of the paper: "Catch-A-Waveform: Learning to Generate Audio from a Single Short Example" (NeurIPS 2021)
★ 189mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4kawesome-datascience. :memo: An awesome Data Science repository to learn and apply for real world problems.
★ 30kchisel-template. A template project for beginning new Chisel work
★ 705yatcpu. Yet another toy CPU.
★ 92yatcpu-docs. Documentation for YatCPU
★ 55opencpop. Opencpop: A High-Quality Open Source Chinese Popular Song Database for Singing Voice Synthesis
★ 237DiffSinger. DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
★ 4.8kphonemizer. Simple text to phones converter for multiple languages
★ 1.6kshennong. A Python toolbox for speech features extraction
★ 166vocoder-benchmark. A repository for benchmarking neural vocoders by their quality and speed.
★ 213chinese-independent-developer. 👩🏿💻👨🏾💻👩🏼💻👨🏽💻👩🏻💻中国独立开发者项目列表 -- 分享大家都在做什么
★ 60ksilero-models. Silero Models: pre-trained text-to-speech models made embarrassingly simple
★ 6kYourTTS. YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone
★ 1.1ktrojan. An unidentifiable mechanism that helps you bypass GFW.
★ 20kserenity. The Serenity Operating System 🐞
★ 34kaioway. AI on the way. An auto deep learning pipe dream. An RDBMS approach to deep learning. Declarative, explainable, scalable, optimizable, easy to deploy, all that good stuff.
★ 1.8kICLR2021-OpenReviewData. Crawl & visualize ICLR papers and reviews.
★ 450CTCDecoder. Connectionist Temporal Classification (CTC) decoding algorithms: best path, beam search, lexicon search, prefix search, and token passing. Implemented in Python.
★ 837python-socketio. Python Socket.IO server and client
★ 4.4ksocket.io. Bidirectional and low-latency communication for every platform
★ 63kwebsocket. Package gorilla/websocket is a fast, well-tested and widely used WebSocket implementation for Go.
★ 25kPatrickStar. PatrickStar enables Larger, Faster, Greener Pretrained Models for NLP and democratizes AI for everyone.
★ 773NUWA. A unified 3D Transformer Pipeline for visual synthesis
★ 2.8kcutlass. CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10kLearn-CUDA-Programming. Learn CUDA Programming, published by Packt
★ 1.3kFlask-SocketIO. Socket.IO integration for Flask applications.
★ 5.5kRAVE. Official implementation of the RAVE model: a Realtime Audio Variational autoEncoder
★ 1.8kTextGAN-PyTorch. TextGAN is a PyTorch framework for Generative Adversarial Networks (GANs) based text generation models.
★ 906redis-py. Redis Python client
★ 14ktransformer-ls. Official PyTorch Implementation of Long-Short Transformer (NeurIPS 2021).
★ 228video-bgm-generation. [ACM MM 2021 Best Paper Award] Video Background Music Generation with Controllable Music Transformer
★ 327flash-linux0.11-talk. 你管这破玩意叫操作系统源码 — 像小说一样品读 Linux 0.11 核心代码
★ 22kStyleSpeech. Official implementation of Meta-StyleSpeech and StyleSpeech
★ 254PaddleSpeech. Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
★ 13kknockknock. 🚪✊Knock Knock: Get notified when your training ends with only two additional lines of code
★ 2.8kdatasets. 🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
★ 22kComputeLibrary. The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologies.
★ 3.2krabbitmq-server. Open source RabbitMQ: core server and tier 1 (built-in) plugins
★ 14k