This is your work, valued

SUP3RMASS1VE

Elite
@SUP3RMASS1VE

Just an amateur finding his way.

Ultimate-TTS-Studio-SUP3R-Edition. 🟒 NVIDIA ONLY – All-in-One TTS App with Kokoro, KittenTTS, Higgs audio, Chatterbox, Fish-Speech, F5 & index-tts & indextts2, Supports Conversation Mode & eBook-to-Audiobook. All features work across all engines in a unified interface except vibe voice which is it's own app panel.

274

VibeVoiceTTS. Frontier Open-Source Text-to-Speech

41

roop-unleashed. A windows pinokio script for roop-unleashed Unsure if it works on other OS

16

Kokoro-TTS-Local-v1.0. Welcome to Kokoro, a high-quality text-to-speech synthesis program powered by deep learning. This tool converts any text into high-fidelity speech in just a few seconds. Simply input text, select a voice, adjust the speed, and enjoy the generated audio.

15

Ultimate-TTS-Studio-SUP3R-Edition-Pinokio. 🟒 NVIDIA ONLY – All-in-One TTS App with Kokoro, KittenTTS, Higgs audio, Chatterbox, Fish-Speech, F5 & index-tts & indextts2, Supports Conversation Mode & eBook-to-Audiobook. All features work across all engines in a unified interface except vibe voice which is it's own app panel.

13

Parakeet-TDT. This is a Gradio web application that uses NVIDIA's Parakeet-TDT-0.6b model for automatic speech recognition with timestamp functionality.

12

chatterbox-SUP3R. SoTA open-source TTS

11

SwarmUI. SwarmUI pinokio script

9

index-tts. An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

8

Nari-Dia-App. Nari Dia is a powerful text-to-speech (TTS) application based on the Dia-1.6B model from Nari Labs. This application allows you to convert text into natural-sounding speech with various customization options.

8

BEN2. BEN2 Background Remover is an AI-powered tool designed to remove backgrounds from images with high precision. It utilizes the BEN_Base model for processing images and is implemented using Gradio for an interactive web-based interface.

8

Re-Size-Image-Outpaint-app. A powerful tool for extending images to different aspect ratios using Stable Diffusion XL.

7

Diffusers-Image-Outpainting. A powerful web application that allows you to expand the boundaries of your images with AI-generated content, creating seamless extensions that match the original image style and context.

7

Deepseek-ai-Janus-Pro-7B. Janus Pro 7B is a powerful multimodal AI model designed for advanced image understanding and text-to-image generation. This project leverages the deepseek-ai/Janus-Pro-7B model, enabling high-quality visual and textual interactions.

7

Realtime-Transcription-app. Real Time Speech Transcription with FastRTC and Local Whisper

6

RuinedFooocus. Forget everything you thought you knew about AI art generation - RuinedFooocus is here to completely reinvent the game! This groundbreaking new image creator combines the best aspects of Stable Diffusion and Midjourney into one seamless, cutting-edge experience. The days of messy installations and manual tweaking are over.

6

VibeVoice-Pinokio. Frontier Open-Source Text-to-Speech (NVIDIA)

6

LiquidAI-LFM2-Audio-1.5B. LFM2-Audio-1.5B is Liquid AI's first end-to-end audio foundation model. Designed with low latency and real time conversation in mind

6

Ovis2-8B-. This project provides a Gradio-based interface for interacting with the Ovis2-8B model. The script allows users to load the model, process image and video inputs, and generate text-based responses using a conversational chatbot.

5

DreamO. DreamO Image Customization. For users with 24GB GPUs, run python app.py --int8 to enable the int8-quantized model. For users with 16GB GPUs, run python app.py --int8 --offload to enable CPU offloading alongside int8 quantization. Note that CPU offload significantly reduces inference speed and should only be enabled when needed

5

IC-Light. More relighting!

5

Kokoro-TTS. OLD VERSION.... Use this one instead Kokoro-TTS-Local-v1.0

5

Image-Upscale. Image Upscale is an AI-powered application designed to enhance and upscale images using advanced techniques like Stable Diffusion and Tile ControlNet. It provides high-quality image enhancement with options for HDR effects and customizable settings.

5

VoxCPM-Text-to-Speech-Pinokio. (NVIDIA)πŸŽ™οΈ VoxCPM 1.5 Text-to-Speech πŸŽ™οΈ

5

Realtime-Transcription. Real Time Speech Transcription with FastRTC and Local Whisper

4

NeuTTS-Air-TTS. State-of-the-art Voice AI has been locked behind web APIs for too long. NeuTTS Air is the world’s first super-realistic, on-device, TTS speech language model with instant voice cloning. Built off a 0.5B

4

YouTube-Downloader. Python

4

Index-TTS-2-Pinokio. (ONLY TESTED ON WINDOWS) IndexTTS2: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

4

Kokoro-TTS-Local. Python

3

Re-Size-Image-Outpaint. (NVIDIA) A powerful tool for extending images to different aspect ratios using Stable Diffusion XL.

3

Spark-TTS. [Windows] and [NVIDIA] Spark-TTS is a high-quality text-to-speech synthesis system that enables voice cloning and speech generation using deep learning models. The system allows users to generate speech with customizable parameters such as pitch, speed, and gender.

3

ChatterBox-Multilingual. SoTA open-source TTS

3

N8N-Pinokio. n8n is a workflow automation platform that gives technical teams the flexibility of code with the speed of no-code. With 400+ integrations, native AI capabilities, and a fair-code license, n8n lets you build powerful automations while maintaining full control over your data and deployments.

3

IC-Light-Ultimate-Studio. This project is an enhanced version of the IC-Light repository, designed for advanced image relighting and enhancement using Stable Diffusion and deep learning techniques.

2

Audio-Transcription-with-Timestamps-parakeet-tdt-0.6b-v2. Python

2

ZipVoice-Pinokio. [NVIDIA} Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

2

AdvancedLivePortraitWebUI. JavaScript

2

Fish-Speech. SOTA Open Source TTS

2

Sana-Sprint-1.6B. Sana, a text-to-image framework that can efficiently generate images up to 4096 Γ— 4096 resolution.

2

MegaTTS-3-Pinokio. MegaTTS 3 is a text-to-speech model trained by ByteDance with exceptional voice cloning capabilities.

2

SD-Next. SD.Next: All-in-one WebUI for AI generative image and video creation

2

UVR5-WebUI. UVR5 UI is an advanced audio separation tool designed to extract individual stems from audio files. This application provides an easy-to-use interface for separating vocals, instruments, and other elements from audio tracks using state-of-the-art machine learning models.

2

FLUX.1-Kontext-Dev-SUP3R. Transform your images with AI-powered editing magic! Upload an image and describe your desired changes - watch the magic happen! All under 20gb VRAM

1

Higgs-Audio-Text-to-Speech-Pinokio. Higgs Audio Text-to-Speech Playground

1

NeuTTS-Air-Pinokio. [NVIDIA] NeuTTS Air is the world’s first super-realistic, on-device, TTS speech language model with instant voice cloning. Built off a 0.5B

1

GPUStack. Simple, scalable AI model deployment on GPU clusters

1

MegaTTS. Python

1

Bagel-DFloat11. [NVIDIA ONLY] Image generation, image editing and free-form manipulation with a VLM

1

Gemini-API-Image-Studio. Python

1

Models.

1