This is your work, valued
Think and do the most important
ComfyUI_IndexTTS. IndexTTS Voice Cloning: Supports two-person dialogue
541Comfyui_HeyGem. HeyGem AI avatar.
265ComfyUI_ACE-Step. ACE-Step: A Step Towards Music Generation Foundation Model
249ComfyUI_MegaTTS3. Lightweight and Efficient, 🎧Ultra High-Quality Voice Cloning, Chinese and English.
211ComfyUI_StepAudioTTS. A Text To Speech node using Step-Audio-TTS in ComfyUI. Can speak, rap, sing, or clone voice.
167ComfyUI_DiffRhythm. Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation. A node for ComfyUI.
153ComfyUI_AudioTools. A ComfyUI node containing multiple audio processing tools.
100ComfyUI_Seed-VC. Seed-VC voice or sing conversion.
66ComfyUI_NotaGen. Symbolic Music Generation, NotaGen node for ComfyUI.
64ComfyUI_Prompt-All-In-One. Prompt Generator for Video, Audio, Image, and Text. A node for ComfyUI. Including Deepseek, Alibaba Cloud Qwen, Google Gemini, and locally selected models, etc.
56ComfyUI_SparkTTS. Using Spark-TTS in Comfyui. Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
51ComfyUI_ASR. 带时间戳、标点符号,自动语音识别。给视频自动添加字幕。
35ComfyUI_KokoroTTS_MW. A Text To Speech node using Kokoro TTS in ComfyUI. Supports 8 languages and 150 voices
33pynini-windows-wheels. pynini-windows-wheels
27ComfyUI_gemmax. ComfyUI Translation Nodes: XiaoMi GemmaX, QuickMT etc.
27ComfyUI_OneButtonPrompt. A node in comfyui for one-click assisted prompt generation (for image and video generation, etc.).
24ComfyUI_PortraitTools. Portrait Tools: Facial detection cropping, alignment, ID photo, etc
21ComfyUI_EraX-WoW-Turbo. Super fast multilingual speech recognition model based on Whisper Large-v3 Turbo. A node for ComfyUI.
14ComfyUI_Dia. Dia TTS model capable of generating ultra-realistic dialogue in one pass. ComfyUI node.
11ComfyUI_OuteTTS. OuteTTS: Multilingual Text-To-Speech, Voice Cloning. A ComfyUI node.
10ComfyUI_SOME. Sing to Midi 🎶
9ComfyUI_DiffRhythm2. DiffRhythm2 歌曲生成。
9ComfyUI_CSM. ComfyUI node of Conversational Speech Model (CSM).
7ComfyUI_parakeet-tdt. parakeet-tdt-0.6b-v2: Automatic speech recognition (ASR) model designed for high-quality English transcription, featuring support for punctuation, capitalization, and accurate timestamp prediction.
4aiart.website. https://aiart.website
2