MB-iSTFT-VITS. Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
469DDSP_Mixture_Model. Demonstration of Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic Sounds
1bittts-demo. HTML
1