This is your work, valued
book-text-to-speech. A book about Text-to-Speech (TTS) in Chinese.
★ 612style-token_tacotron2. style token with tacotron2
★ 62blog. personal blog
★ 18speech_emotion. Detect emotion from audio
★ 14tpse_tacotron2. TPSE-GST Tacotron2
★ 14tacotron2. Python
★ 12public_praise_prediction_yunyi. 2018云移杯景区口碑评价分值预测 7/1186
★ 11LLM-paper-daily. Automatically Update LLM Papers Daily using Github Actions. Ref: https://github.com/Vincentqyw/cv-arxiv-daily
★ 10deep_learning_practice. This is my study & practice repository of some deep learning models which I am interested in
★ 7kafka_spark_streaming. A very simple example of using streaming data by kafka & spark streaming & mongodb & bokeh
★ 6emotion_recognization. recognize emotion from text and speech
★ 4tacotron2decoder. pretrain tacotron2 decoder
★ 4SpeechQuality. Python
★ 4HTS-Project. HTS-demo project with blank data, expecially extended for Mandarin Chinese SPSS system.
★ 2crnn_sound_classification. Python
★ 2differential_evolution. Differential Evolution algorithm
★ 2Tacotron-2. DeepMind's Tacotron-2 Tensorflow implementation
★ 2AudioProcessBasis. C
★ 1algorithm_practise. This is the code for recruitment :(
★ 1deeplearningbook-chinese. Deep Learning Book Chinese Translation
★ 1Multilingual_Text_to_Speech. An implementation of Tacotron 2 that supports multilingual experiments with parameter-sharing, code-switching, and voice cloning.
★ 1Shaper_SmartHome. Shaper smart-home project
★ 1awesome-code-snippets. Some useful and clever code🍺
★ 1listen-and-rate. A lightweight, self-hostable tool for conducting subjective listening tests in a browser.
★ 29DuplexOmni. Python
★ 64asr-hotword. 最棒的的ASR后处理热词方案,基于音素编辑距离,实现热词替换。
★ 45tts-bench. Speed and samples benchmark: for all types of text to speech (TTS) models on Windows/Linux/Mac.
★ 262forced-aligners-bench. Benchmark forced alignment models (Qwen3-FA, WhisperX, Seamless UnitY2) on FLEURS via implicit WER on cropped audio.
★ 11CrossLingual-TTS-Lab. A small benchmark harness for testing voice-language disentanglement in multilingual text-to-speech systems.
★ 8modern-gpu-programming-for-mlsys. A tutorial on modern GPU programming for machine learning systems
★ 1.1kopen-audio-opd. Industrial audio online policy distillation (OPD) training stack for ASR and TTS, distilling compact audio models from stronger teacher models.
★ 634ai-engineering-from-scratch. Learn it. Build it. Ship it for others.
★ 45kvllm-omni. A framework for efficient model inference with omni-modality models
★ 5.7kawesome-mcp-servers. A collection of MCP servers.
★ 92kminimind-o. 🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!
★ 2.2kAI-Infra-Auto-Driven-SKILLS. Python
★ 707hermes-agent. The agent that grows with you
★ 223kRelax. An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 553scaled-echo-tts. Scaled diffusion transformer for text-to-speech synthesis (DiT + T5Gemma2 conditioning, TorchTitan & Megatron backends, tested up to 1024 GPUs)
★ 24claude-code. Claude Code 源码文档解析
★ 68InfraTech. 分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
★ 3.2kSoulX-Duplug. Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.
★ 281Resonate. [INTERSPEECH 2026] Pre-training, SFT, DPO and GRPO for Text-to-Audio Generation
★ 48FireRedVAD. A SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD
★ 477compute-wer. Compute WER and SER for speech recognition evaluation
★ 28Qwen3-TTS. Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
★ 13kFinetune_Nemo_ASR. Finetune Nemo parakeet ASR model with new language (support 8 bit optimizer). Experimental birwkv-fastconformer TDT for long-form ASR(8.5 hours in single pass).
★ 26senko. Very fast, accurate speaker diarization
★ 287nano-vllm. Nano vLLM
★ 15knano-whisper. A demo-level low-latency, high-throughput inference engine for whisper
★ 20werpy. 🐍📦 Ultra-fast Python package for calculating and analyzing the Word Error Rate (WER). Built for the scalable evaluation of speech and transcription accuracy.
★ 29Ming-UniAudio. Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
★ 452error-align. Text-to-text alignment algorithm for speech recognition error analysis.
★ 32MiniMind-in-Depth. 轻量级大语言模型MiniMind的源码解读,包含tokenizer、RoPE、MoE、KV Cache、pretraining、SFT、LoRA、DPO等完整流程
★ 1.1kjson_repair. Repair malformed JSON from LLMs, APIs, logs, and user input in Python.
★ 5.1kVibeVoice. Open-Source Frontier Voice AI
★ 52kSimulStreaming. Python
★ 645WhisperLiveKit. Simultaneous speech-to-text models
★ 11kRLFromScratch. Python
★ 644MOSS-TTSD. MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.
★ 1.4khiggs-audio. Text-audio foundation model from Boson AI
★ 8.3kzhvoice. Chinese voice corpus. 中文语音语料,语音更加清晰自然,包含8个开源数据集,3200个说话人,900小时语音,1300万字。
★ 749UTMOSv2. UTokyo-SaruLab MOS Prediction System
★ 357Diffusion-LLM-Papers. A Collection of Papers on Diffusion Language Models
★ 181ConversationTTS. Python
★ 101Spark-TTS. Spark-TTS Inference Code
★ 11kten-vad. Voice Activity Detector (VAD) : low-latency, high-performance and lightweight
★ 2.2kCTCDataset. 中文文本纠错数据集汇总
★ 47Muyan-TTS. Python
★ 481Full-Duplex-Bench. A Benchmark for Evaluating Turn-Taking and Overlap Handling in Full-Duplex Spoken Dialogue Models
★ 247Kimi-Audio. Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
★ 4.7kWhisper-Finetune. Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
★ 318Awesome-LLM-Long-Context-Modeling. 📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
★ 2.1kAwesome-Simultaneous-Translation. Paper list of simultaneous translation / streaming translation, including text-to-text machine translation and speech-to-text translation.
★ 579MultiMed-ST. [EMNLP 2025] MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
★ 148VoxBox. A large-scale speech corpus introduced in Spark-TTS, built from diverse open-source datasets for training text-to-speech (TTS) systems.
★ 115codellm-data-preprocess-pipeline. 代码大模型 预训练&微调&DPO 数据处理 业界处理pipeline sota
★ 53opc_data_filtering. Heuristic filtering framework for RefineCode
★ 87Orpheus-TTS. Towards Human-Sounding Speech
★ 6.3kLLMDataHub. A quick guide (especially) for trending instruction finetuning datasets
★ 3.4kllm-datasets. Curated list of datasets and tools for post-training.
★ 4.7kBaichuan-Audio. Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
★ 223Awesome-LLM-Reasoning. From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
★ 3.7krllm. Democratizing Reinforcement Learning for LLMs
★ 5.7kAIInfra. AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。
★ 7.8kaudiobox-aesthetics. Unified automatic quality assessment for speech, music, and sound.
★ 748TinyZero. Minimal reproduction of DeepSeek R1-Zero
★ 13ks1. s1: Simple test-time scaling
★ 6.7kMiniMax-01. The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
★ 3.4kopen-r1. Fully open reproduction of DeepSeek-R1
★ 26kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kawesome-LLM-resources. 🧑🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.
★ 8.8kMiniCPM-V. A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
★ 26kminimind-v. 👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!
★ 8.4kX-Codec-2.0. Codec for paper: LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
★ 362Papers-in-100-Lines-of-Code. Implementation of papers in 100 lines of code.
★ 2.8kawesome-controllable-speech-synthesis. This is an evolving repo for the paper "Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey".
★ 276CrisperWhisper. Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection
★ 1kllm-course. Jupyter Notebook
★ 883llm_interview_note. 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
★ 15kTTS-arxiv-daily. Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)
★ 663Neural-Codec-and-Speech-Language-Models. Awesome Neural Codec Models, Text-to-Speech Synthesizers & Speech Language Models
★ 246GLM-4-Voice. GLM-4-Voice | 端到端中英语音对话模型
★ 3.2ktorchtitan. A PyTorch native platform for training generative AI models
★ 5.6ktorchtune. PyTorch native post-training library
★ 5.8klingua. Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
★ 4.8kDM-Codec. Source code for the EMNLP 2025 paper “DM-Codec: Distilling Multimodal Representations for Speech Tokenization”
★ 57promptttspp. PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions
★ 86scoreq. SCOREQ: Speech COntrastive REgression for Quality Assessment (NeurIPS 2024)
★ 115whisper-timestamped. Multilingual Automatic Speech Recognition with word-level timestamps and confidence
★ 2.8kJEmoji. Java Emoji (JEmoji) is a lightweight, fast and auto generated emoji library for Java with the purpose to improve and ease working with emojis
★ 115Open-O1. Python
★ 1.3kGTSinger. Dataset and code of GTSinger(NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
★ 439efts2code. source code of EfficientTTS 2
★ 21ms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kpytorch_stoi. STOI loss function in PyTorch
★ 107ichigo. Local realtime voice AI
★ 2.5kmosel. Collection of Open Source Speech Data
★ 166FireRedTTS. An Open-Sourced LLM-empowered Foundation TTS System
★ 909LLM-Travel. 欢迎来到 "LLM-travel" 仓库!探索大语言模型(LLM)的奥秘 🚀。致力于深入理解、探讨以及实现与大模型相关的各种技术、原理和应用。
★ 385RSTnet. Real-time Speech-Text Foundation Model Toolkit (wip)
★ 255moshi. Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
★ 11ksuper-monotonic-align. Python
★ 173so-large-lm. 大模型基础: 一文了解大模型基础知识
★ 7.5kBayLing-Speech. LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
★ 3.1kllm-action. 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
★ 25kmini-omni. open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
★ 3.6kAwesome-Unified-Multimodal-Models. 📖 This is a repository for organizing papers, codes and other resources related to unified multimodal models.
★ 829audio-ai-hub. The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
★ 949MooER. MooER: Moore-threads Open Omni model for speech-to-speech intERaction. MooER-omni includes a series of end-to-end speech interaction models along with training and inference code, covering but not limited to end-to-end speech interaction, end-to-end speech translation and speech recognition.
★ 219ltu. Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
★ 478VITA. ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
★ 2.5kwesr. We Speech Transcript based on LLM, in 300 lines of code.
★ 182audio-preprocess. Preprocess Audio for training
★ 399ultravox. A fast multimodal LLM for real-time voice
★ 4.5kPDF-Extract-Kit. A Comprehensive Toolkit for High-Quality PDF Content Extraction
★ 9.8kNoresqa. This github repo is for Neurips 2021 and Interspeech 2022 papers on Non-Matching Reference based estimation of speech quality assessment.
★ 104Qwen2-Audio. The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 2.1kopen_flamingo. An open-source framework for training large multimodal models.
★ 4.1kaudio-flamingo. PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
★ 1.2kAnyGPT. Code for "AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling"
★ 880awesome-RLHF. A curated list of reinforcement learning with human feedback resources (continually updated)
★ 4.4kDPO. dpo算法实现
★ 53Llama-Chinese. Llama中文社区,实时汇总最新Llama学习资料,构建最好的中文Llama大模型开源生态,完全开源可商用
★ 15kAwesome-ChatTTS. 官方推荐的 ChatTTS 资源汇总项目,整理了全网相关资源和常见问题 || Officially recommended ChatTTS resource collection project
★ 1.9kControlSpeech. [ACL 2025 Main] ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec
★ 276LLMBox. A comprehensive library for implementing LLMs, including a unified training pipeline and comprehensive model evaluation.
★ 847ChatTTS. A generative speech model for daily dialogue.
★ 40kYulan-GARDEN. Official Repository for SIGIR2024 Demo Paper "An Integrated Data Processing Framework for Pretraining Foundation Models"
★ 88SLAM-LLM. A Framework for Speech, Language, Audio, Music Processing with Large Language Model
★ 1.1kllama3-from-scratch. llama3 implementation one matrix multiplication at a time
★ 15kLLMBook-zh.github.io. 《大语言模型》作者:赵鑫,李军毅,周昆,唐天一,文继荣
★ 4.5kchinese-chatbot-corpus. 中文公开聊天语料库
★ 4.2kReazonSpeech. Massive open Japanese speech corpus
★ 390FunClip. FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
★ 6.1kpython-patterns. A collection of design patterns/idioms in Python
★ 43kdeep-diacritization. Official Repository of the Deep Diacritization Paper
★ 17llama3-Chinese-chat. Llama3-中文后训练版
★ 4.1kllama3. The official Meta Llama 3 GitHub site
★ 29kdataspeech. Python
★ 400speech-trident. Awesome speech/audio LLMs, representation learning, and codec models
★ 1.2kABigSurveyOfLLMs. A collection of 150+ surveys on LLMs
★ 351StableTTS. Next-generation TTS model using flow-matching and DiT, inspired by Stable Diffusion 3
★ 438baby-llama2-chinese. 用于从头预训练+SFT一个小参数量的中文LLaMa2的仓库;24G单卡即可运行得到一个具备简单中文问答能力的chat-llama2.
★ 2.9klm-evaluation-harness. A framework for few-shot evaluation of language models.
★ 13kBigListOfPodcasts. A list of podcast URLs scraped from the Apple podcast database in late 2021, including a script for downloading those podcasts.
★ 44VoiceCraft. Zero-Shot Speech Editing and Text-to-Speech in the Wild
★ 8.5kOpenPhonemizer. An espeak-compatible, permissively-licensed IPA phonemizer (G2P) based on DeepPhonemizer. Usable as a drop-in replacement for espeak's GPL phonemizer.
★ 112LLM4Decompile. Reverse Engineering: Decompiling Binary Code with Large Language Models
★ 6.8kOpen-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kComfyUI-Workflows-ZHO. 我的 ComfyUI 工作流合集 | My ComfyUI workflows collection
★ 7.7kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kcontent_aware_mos_prediction. A suite of ML models that predict the quality of synthetic speech generated by text-to-speech (TTS) models. Estimates the mean opinion score (MOS) of synthetic audio snippets
★ 7AudioEditingCode. Python
★ 195Sequoia. scalable and robust tree-based speculative decoding algorithm
★ 377awesome-audio-plaza. Daily tracking of awesome audio papers, including music generation, zero-shot tts, asr, audio generation
★ 410MeloTTS. High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
★ 7.6kZMM-TTS. ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
★ 185awesome-chatgpt-dataset. Unlock the Power of LLM: Explore These Datasets to Train Your Own ChatGPT!
★ 765awesome-foundation-and-multimodal-models. 👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
★ 637speech-dataset-generator. 🔊 Create labeled datasets, enhance audio quality, identify speakers, support diverse dataset types. 🎧👥📊 Advanced audio processing.
★ 262Languagecodec. [ACL 2025 Oral] Language-Codec: Reducing the Gaps Between Discrete Codec Representation and Speech Language Models
★ 208Emu. Emu Series: Generative Multimodal Models from BAAI
★ 1.8kgemma_pytorch. The official PyTorch implementation of Google's Gemma models
★ 5.7ksnac. Multi-Scale Neural Audio Codec (SNAC) compresses audio into discrete codes at a low bitrate
★ 774torch-audiomentations. Fast audio data augmentation in PyTorch. Inspired by audiomentations. Useful for deep learning.
★ 1.2kstable-audio-tools. Generative models for conditional audio generation
★ 3.8kagc. Audiogen Codec
★ 146Codec-SUPERB. Audio Codec Speech processing Universal PERformance Benchmark
★ 309minbpe. Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.
★ 11kGenshin_Datasets. Genshin Datasets For SVC/SVS/TTS
★ 735aimoneyhunter. ai副业赚钱大集合,教你如何利用ai做一些副业项目,赚取更多额外收益。The Ultimate Guide to Making Money with AI Side Hustles: Learn how to leverage AI for some cool side gigs and rake in some extra cash. Check out the English version for more insights.
★ 18kstable-audio-metrics. Metrics for evaluating music and audio generative models – with a focus on long-form, full-band, and stereo generations.
★ 300HiFiPLN. Multispeaker Community Vocoder Model for DiffSinger
★ 39supervoice-voicebox. VoiceBox neural network implementation
★ 110FlagEmbedding. Retrieval and Retrieval-augmented LLMs
★ 12kNeural-Transducers-for-Two-Stage-Text-to-Speech-via-Semantic-Token-Prediction. Unofficial pytorch reproduction for the paper "Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction" (arXiv:2401.01498)
★ 60natural_voice_assistant. Python
★ 499metavoice-src. Foundational model for human-like, expressive TTS
★ 4.2kPhishingBook. 红蓝对抗:钓鱼演练资源汇总&备忘录
★ 1.2kvallt.
★ 36g2p-zh-en. Chinese and English Bilinguish G2P
★ 22youtube-8m-videos-downloader. Download videos from YouTube-8M dataset for testing
★ 6spiking-fullsubnet. Official repository of Spiking-FullSubNet, the Intel N-DNS Challenge Algorithmic Track Winner.
★ 142vocal-separate. an extremely simple tool for separating vocals and background music, completely localized for web operation, using 2stems/4stems/5stems models 这是一个极简的人声和背景音乐分离工具,本地化网页操作,无需连接外网
★ 2kASR_TOOLS_SenseVoice_WebUI. Bert-vits2转写和标注独立整合Webui,整合阿里FunAsr,必剪Asr以及Whisper大模型
★ 182GPT-SoVITS. 1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 60kopen-tts-tracker.
★ 1.1kSvarah. Swarah: Indian-English speech dataset collected across the country
★ 38Anim400K. Anim-400K: A dataset designed from the ground up for automated dubbing of video
★ 118onnx-tool. A tool for parsing, editing, optimizing, and profiling ONNX models.
★ 491nansypp. Unofficial implementation of NANSY++ in Pytorch Lightning
★ 50tts-frontend-dataset. TTS FrontEnd DataSet: Polyphone / Prosody / TextNormalization
★ 104pycantonese. Cantonese Linguistics and NLP
★ 413pheme. Python
★ 258megatts2. Unoffical implementation of Megatts2
★ 285vocal-remover. Vocal Remover using Deep Neural Networks
★ 1.8kColaboratory-Notebook-for-Ultimate-Vocal-Remover. Colaboratory Notebook for Ultimate Vocal Remover
★ 100TikTokDownloader. TikTok 发布/喜欢/合辑/直播/视频/图集/音乐;抖音发布/喜欢/收藏/收藏夹/视频/图集/实况/直播/音乐/合集/评论/账号/搜索/热榜数据采集工具/下载工具
★ 15kOpenVoice. Instant voice cloning by MIT and MyShell. Audio foundation model.
★ 37kWhisperS2T. An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
★ 577insanely-fast-whisper. Incredibly fast Whisper-large-v3
★ 1.9kMahaTTS. Python
★ 275