lecture_dtw_notebook. Jupyter Notebook
80bayes_gmm. Bayesian Gaussian mixture models in Python.
64data414. Data Analytics 414
57vqwordseg. Unsupervised phone and word segmentation using dynamic programming on self-supervised VQ features.
39linearvc. Voice conversion with just linear regression.
37autoencoders_mnist. Different types of autoencoders illustrated on MNIST using TensorFlow.
36speech_dtw. Dynamic time warping (DTW) functions for specifically speech alignment.
30nlp817. Natural Language Processing 817
24recipe_swbd_wordembeds. Shell
22couscous. Siamese neural networks for representation learning using Theano.
20eskmeans. Embedded segmental K-means (ES-KMeans) in Python.
14speech_correspondence. Correspondence and autoencoder neural network training for speech using Pylearn2.
14segmentalist. Unsupervised word segmentation and clustering of speech
13globalphone_awe. Multilingual acoustic word embedding approaches applied and evaluated on GlobalPhone data.
11suzerospeech2019. Stellenbosch University ZeroSpeech 2019 System
10recipe_zs2017_track2. Recipe for applying the embedded segmental k-means model to the ZeroSpeech2017 Track 2 challenge.
9recipe_bucktsong_awe_py3. Unsupervised acoustic word embeddings evaluated on Buckeye English and NCHLT Xitsonga data in Python 3.
9bucktsong_segmentalist. Unsupervised segmentation and clustering of Buckeye English and NCHLT Xitsonga corpora.
9yfacc. YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding
8semantic_flickraudio. A data set of semantic keyword spotting labels for the Flickr Audio Captions Corpus.
7recipe_vision_speech_flickr. Using computer vision to ground unlabelled speech.
7stellenbosch_ee_report_template. A LaTeX template for reports following the guidelines of the E&E department at Stellenbosch University.
6recipe_bucktsong_awe. Unsupervised acoustic word embeddings evaluated on Buckeye English and NCHLT Xitsonga data in Python 2.7.
6recipe_semantic_flickraudio. Semantic speech retrieval with a visually grounded model of untranscribed speech.
4bucktsong_eskmeans. Unsupervised segmentation and clustering of the Buckeye English and NCHLT Xitsonga datasets using the ES-KMeans algorithm.
3kamperh.github.io. SCSS
3dpdp_aernn. Duration-penalized dynamic programming (DPDP) autoencoding recurrent neural network (AE-RNN) in Python.
2ss414. Systems and Signals 414
2flickr_semantic_qbe_eval. Evaluation code for semantic QbE on the Flickr8k Audio Captions Corpus
1teaching_portfolio. Teaching Portfolio
1