Research Scientist @ Facebook AI Research (FAIR). Former PhD Student @ MIT Spoken Language Systems Group
FactorizedHierarchicalVAE. This repository contains the code to reproduce the core results from the paper "Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data"
155ScalableFHVAE. This repository contains the code to reproduce the core results from the paper "Scalable Factorized Hierarchical Variational Autoencoders"
53SpeechVAE. This repository contains the code to reproduce the core results from the paper "Learning Latent Representations for Speech Generation and Transformation".
52ResDAVEnet-VQ. Official codes for the paper "Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech"
28ReVISE. ReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech Enhancement
14PGLSTM_ASR. This repo contains codes to reproduce the core results of "A Prioritized Grid Long Short-Term Memory RNN for Speech Recognition"
3semi-supervised-pytorch. Implementations of different VAE-based semi-supervised and generative models in PyTorch
3tensorflow-wavenet. A TensorFlow implementation of DeepMind's WaveNet paper
2ZeroSpeech2019_RLE_eval. ZeroSpeech 2019 evaluation with run-length encoding (RLE), metrics reported in ResDAVEnet-VQ.
1tacotron2_dev. Jupyter Notebook
1wavenet_vocoder. WaveNet vocoder
1