Audio/AI researcher @ Sony Research (Sony AI), working on audio generative models, beer lover
friendly-stable-audio-tools. Refactored / updated version of `stable-audio-tools` which is an open-source code for audio/music generative models originally by Stability AI.
218floss-torch. PyTorch implementation of "Source Separation by Flow Matching (FLOSS)" by Google DeepMind
97modified-shortcut-models-pytorch. PyTorch implementation of Shortcut Models [Frans, 2025] with little modification
71Open-Miipher-2. PyTorch implementation of Miipher-2 [2025] which is a speech restoration model by Google DeepMind
70wavefit-pytorch. PyTorch implementation of WaveFit [2022, Google] which is one of SOTA lightweight/fast speech vocoders.
70minimal-musicgen-for-developers. [PyTorch] Minimal codebase for MusicGen models
63world-class. A C++ library of "World" - A high-quality speech analysis, manipulation and synthesis system -
60minimal-sqvae. A minimal Pytorch Implementation of Stochastically Quantized Variational AutoEncoder (SQ-VAE) by Sony
33Swin-Transformer-1d. PyTorch implementation of Swin Transformer for 1-dimensional data
19minimal-san. A minimal Pytorch Implementation of a Slicing Adversarial Network by Sony
7online-rpca. Jupyter Notebook
5abci-code-sample. Python
1