:fist_raised: MIT Postdoc :fist_raised: NTU Ph.D. :fist_raised: former Research Scientist Intern @ Meta Reality labs
speech-trident. Awesome speech/audio LLMs, representation learning, and codec models
1.2kML2021-Spring. **Official** 李宏毅 (Hung-yi Lee) 機器學習 Machine Learning 2021 Spring
1kSpeech-Prompts-Adapters. This Repository surveys the paper focusing on Prompting and Adapters for Speech Processing.
113SpeechPrompt. **Interspeech 2022** 《SpeechPrompt: An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks》Speech processing with prompting paradigm
102FastSpeech2. Multi-Speaker Pytorch FastSpeech2: Fast and High-Quality End-to-End Text to Speech :fist:
99SpeechPrompt-v2. 《SpeechPrompt v2: Prompt Tuning for Speech Classification Tasks》Speech processing with prompting paradigm
81SpeechGen. 《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》
77Taiwanese-Whisper. fine-tune Whipser model for Taiwanese speech recognition
37Spoken-Dialogue-Model-Survey. A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation Pattern
31Taiwanese-Speech-Synthesis. Taiwanese Speech Synthesis with Tacotron2
26AudioCodec-Hub. AudioCodec-Hub is a Python library for encoding and decoding audio data, supporting various neural audio codec models
25RobustVC. **ICASSP 2022** 《Toward Degradation-Robust Voice Conversion》Using speech enhancement and end-to-end denoising training to improve degradation / adversarial robustness of VC models.
24Taiwanese-Translation. Taiwanese Translation with BERT based model and RNN. Collection of Taiwanese text corpus
13TaiwaneseTTS. Python
10FlappyBird. :fire: Super Flappy Bird in p5.js
10Game-Time-Benchmark. Game-Time: Evaluating Temporal Dynamics in Full-Duplex Spoken Language Models
10moth. 虫我研所 Moth Institute 新一代設計展 https://ga642381.github.io/moth
8Kai-Wei-Chang-Talks. A repository sharing slides of the talks I gave
7FinanceWeb. JavaScript
6rocling2025-proceedings. Python
5seamless_communication_emo. Foundational Models for State-of-the-Art Speech and Text Translation
3S2VC. Python
3neurips2021-sas-react. JavaScript
3CA2021-Final. Jupyter Notebook
3Deep-Q-learning. Playing Atari game (breakout) with deep reinforcement learning
3Taiwan-Stock-Crawler. Taiwan-Stock-Crawler, crawling data from TWSE
3s3prl. Self-Supervised Speech Pre-training and Representation Learning Toolkit.
2TensorFlowTTS. :stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, Korean, Chinese and Easy to adapt for other languages)
2ELK-Stack.
1vision. Datasets, Transforms and Models specific to Computer Vision
1Linguistics-111. Jupyter Notebook
1awesome-llm-role-playing-with-persona. Awesome-llm-role-playing-with-persona: a curated list of resources for large language models for role-playing with assigned personas
1LSGAN. LSGAN for image generation on anime dataset
1AudioDec. An Open-source Streaming High-fidelity Neural Audio Codec
1speech_quality. Jupyter Notebook
1Codec-SUPERB. Audio Codec Speech processing Universal PERformance Benchmark
1speech-language-model. A collection of papers related to speech language models
1