Focusing on multimodal synthesis (speech/audio/music), speech translation, and audio editing.
OmniAudio. [ICML 2025] PyTorch Implementation of "OmniAudio: Generating Spatial Audio from 360-Degree Video"
375AudioLCM. PyTorch Implementation of [AudioLCM]: a efficient and high-quality text-to-audio generation with latent consistency model.
13ViT-TTS. PyTorch Implementation of ViT-TTS (EMNLP'23)
11FlashAudio. PyTorch Implementation of FlashAudio with Rectified Flow Models in Text-to-Audio Generation
5Persona-Dialogue-Generation. a repository for persona multi-turn dialogue
5Sphere360. A 360-degree video dataset designed for 360-degree video-to-spatial audio generation.
4MEDIC. PyTorch Implementation of MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
4