This is your work, valued
Just an amateur finding his way.
Ultimate-TTS-Studio-SUP3R-Edition. π’ NVIDIA ONLY β All-in-One TTS App with Kokoro, KittenTTS, Higgs audio, Chatterbox, Fish-Speech, F5 & index-tts & indextts2, Supports Conversation Mode & eBook-to-Audiobook. All features work across all engines in a unified interface except vibe voice which is it's own app panel.
274VibeVoiceTTS. Frontier Open-Source Text-to-Speech
41roop-unleashed. A windows pinokio script for roop-unleashed Unsure if it works on other OS
16Kokoro-TTS-Local-v1.0. Welcome to Kokoro, a high-quality text-to-speech synthesis program powered by deep learning. This tool converts any text into high-fidelity speech in just a few seconds. Simply input text, select a voice, adjust the speed, and enjoy the generated audio.
15Ultimate-TTS-Studio-SUP3R-Edition-Pinokio. π’ NVIDIA ONLY β All-in-One TTS App with Kokoro, KittenTTS, Higgs audio, Chatterbox, Fish-Speech, F5 & index-tts & indextts2, Supports Conversation Mode & eBook-to-Audiobook. All features work across all engines in a unified interface except vibe voice which is it's own app panel.
13Parakeet-TDT. This is a Gradio web application that uses NVIDIA's Parakeet-TDT-0.6b model for automatic speech recognition with timestamp functionality.
12chatterbox-SUP3R. SoTA open-source TTS
11SwarmUI. SwarmUI pinokio script
9index-tts. An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
8Nari-Dia-App. Nari Dia is a powerful text-to-speech (TTS) application based on the Dia-1.6B model from Nari Labs. This application allows you to convert text into natural-sounding speech with various customization options.
8BEN2. BEN2 Background Remover is an AI-powered tool designed to remove backgrounds from images with high precision. It utilizes the BEN_Base model for processing images and is implemented using Gradio for an interactive web-based interface.
8Re-Size-Image-Outpaint-app. A powerful tool for extending images to different aspect ratios using Stable Diffusion XL.
7Diffusers-Image-Outpainting. A powerful web application that allows you to expand the boundaries of your images with AI-generated content, creating seamless extensions that match the original image style and context.
7Deepseek-ai-Janus-Pro-7B. Janus Pro 7B is a powerful multimodal AI model designed for advanced image understanding and text-to-image generation. This project leverages the deepseek-ai/Janus-Pro-7B model, enabling high-quality visual and textual interactions.
7Realtime-Transcription-app. Real Time Speech Transcription with FastRTC and Local Whisper
6RuinedFooocus. Forget everything you thought you knew about AI art generation - RuinedFooocus is here to completely reinvent the game! This groundbreaking new image creator combines the best aspects of Stable Diffusion and Midjourney into one seamless, cutting-edge experience. The days of messy installations and manual tweaking are over.
6VibeVoice-Pinokio. Frontier Open-Source Text-to-Speech (NVIDIA)
6LiquidAI-LFM2-Audio-1.5B. LFM2-Audio-1.5B is Liquid AI's first end-to-end audio foundation model. Designed with low latency and real time conversation in mind
6Ovis2-8B-. This project provides a Gradio-based interface for interacting with the Ovis2-8B model. The script allows users to load the model, process image and video inputs, and generate text-based responses using a conversational chatbot.
5DreamO. DreamO Image Customization. For users with 24GB GPUs, run python app.py --int8 to enable the int8-quantized model. For users with 16GB GPUs, run python app.py --int8 --offload to enable CPU offloading alongside int8 quantization. Note that CPU offload significantly reduces inference speed and should only be enabled when needed
5IC-Light. More relighting!
5Kokoro-TTS. OLD VERSION.... Use this one instead Kokoro-TTS-Local-v1.0
5Image-Upscale. Image Upscale is an AI-powered application designed to enhance and upscale images using advanced techniques like Stable Diffusion and Tile ControlNet. It provides high-quality image enhancement with options for HDR effects and customizable settings.
5VoxCPM-Text-to-Speech-Pinokio. (NVIDIA)ποΈ VoxCPM 1.5 Text-to-Speech ποΈ
5Realtime-Transcription. Real Time Speech Transcription with FastRTC and Local Whisper
4NeuTTS-Air-TTS. State-of-the-art Voice AI has been locked behind web APIs for too long. NeuTTS Air is the worldβs first super-realistic, on-device, TTS speech language model with instant voice cloning. Built off a 0.5B
4YouTube-Downloader. Python
4Index-TTS-2-Pinokio. (ONLY TESTED ON WINDOWS) IndexTTS2: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
4Kokoro-TTS-Local. Python
3Re-Size-Image-Outpaint. (NVIDIA) A powerful tool for extending images to different aspect ratios using Stable Diffusion XL.
3Spark-TTS. [Windows] and [NVIDIA] Spark-TTS is a high-quality text-to-speech synthesis system that enables voice cloning and speech generation using deep learning models. The system allows users to generate speech with customizable parameters such as pitch, speed, and gender.
3ChatterBox-Multilingual. SoTA open-source TTS
3N8N-Pinokio. n8n is a workflow automation platform that gives technical teams the flexibility of code with the speed of no-code. With 400+ integrations, native AI capabilities, and a fair-code license, n8n lets you build powerful automations while maintaining full control over your data and deployments.
3IC-Light-Ultimate-Studio. This project is an enhanced version of the IC-Light repository, designed for advanced image relighting and enhancement using Stable Diffusion and deep learning techniques.
2Audio-Transcription-with-Timestamps-parakeet-tdt-0.6b-v2. Python
2ZipVoice-Pinokio. [NVIDIA} Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
2AdvancedLivePortraitWebUI. JavaScript
2Fish-Speech. SOTA Open Source TTS
2Sana-Sprint-1.6B. Sana, a text-to-image framework that can efficiently generate images up to 4096 Γ 4096 resolution.
2MegaTTS-3-Pinokio. MegaTTS 3 is a text-to-speech model trained by ByteDance with exceptional voice cloning capabilities.
2SD-Next. SD.Next: All-in-one WebUI for AI generative image and video creation
2UVR5-WebUI. UVR5 UI is an advanced audio separation tool designed to extract individual stems from audio files. This application provides an easy-to-use interface for separating vocals, instruments, and other elements from audio tracks using state-of-the-art machine learning models.
2FLUX.1-Kontext-Dev-SUP3R. Transform your images with AI-powered editing magic! Upload an image and describe your desired changes - watch the magic happen! All under 20gb VRAM
1Higgs-Audio-Text-to-Speech-Pinokio. Higgs Audio Text-to-Speech Playground
1NeuTTS-Air-Pinokio. [NVIDIA] NeuTTS Air is the worldβs first super-realistic, on-device, TTS speech language model with instant voice cloning. Built off a 0.5B
1GPUStack. Simple, scalable AI model deployment on GPU clusters
1MegaTTS. Python
1Bagel-DFloat11. [NVIDIA ONLY] Image generation, image editing and free-form manipulation with a VLM
1Gemini-API-Image-Studio. Python
1Models.
1