Cambridge, USA

Kai-Wei Chang (張凱爲)

Elite
@ga642381

:fist_raised: MIT Postdoc :fist_raised: NTU Ph.D. :fist_raised: former Research Scientist Intern @ Meta Reality labs

speech-trident. Awesome speech/audio LLMs, representation learning, and codec models

1.2k

ML2021-Spring. **Official** 李宏毅 (Hung-yi Lee) 機器學習 Machine Learning 2021 Spring

1k

Speech-Prompts-Adapters. This Repository surveys the paper focusing on Prompting and Adapters for Speech Processing.

113

SpeechPrompt. **Interspeech 2022** 《SpeechPrompt: An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks》Speech processing with prompting paradigm

102

FastSpeech2. Multi-Speaker Pytorch FastSpeech2: Fast and High-Quality End-to-End Text to Speech :fist:

99

SpeechPrompt-v2. 《SpeechPrompt v2: Prompt Tuning for Speech Classification Tasks》Speech processing with prompting paradigm

81

SpeechGen. 《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》

77

Taiwanese-Whisper. fine-tune Whipser model for Taiwanese speech recognition

37

Spoken-Dialogue-Model-Survey. A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation Pattern

31

Taiwanese-Speech-Synthesis. Taiwanese Speech Synthesis with Tacotron2

26

AudioCodec-Hub. AudioCodec-Hub is a Python library for encoding and decoding audio data, supporting various neural audio codec models

25

RobustVC. **ICASSP 2022** 《Toward Degradation-Robust Voice Conversion》Using speech enhancement and end-to-end denoising training to improve degradation / adversarial robustness of VC models.

24

Taiwanese-Translation. Taiwanese Translation with BERT based model and RNN. Collection of Taiwanese text corpus

13

TaiwaneseTTS. Python

10

FlappyBird. :fire: Super Flappy Bird in p5.js

10

Game-Time-Benchmark. Game-Time: Evaluating Temporal Dynamics in Full-Duplex Spoken Language Models

10

moth. 虫我研所 Moth Institute 新一代設計展 https://ga642381.github.io/moth

8

Kai-Wei-Chang-Talks. A repository sharing slides of the talks I gave

7

FinanceWeb. JavaScript

6

rocling2025-proceedings. Python

5

seamless_communication_emo. Foundational Models for State-of-the-Art Speech and Text Translation

3

S2VC. Python

3

neurips2021-sas-react. JavaScript

3

CA2021-Final. Jupyter Notebook

3

Deep-Q-learning. Playing Atari game (breakout) with deep reinforcement learning

3

Taiwan-Stock-Crawler. Taiwan-Stock-Crawler, crawling data from TWSE

3

s3prl. Self-Supervised Speech Pre-training and Representation Learning Toolkit.

2

TensorFlowTTS. :stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, Korean, Chinese and Easy to adapt for other languages)

2

ELK-Stack.

1

vision. Datasets, Transforms and Models specific to Computer Vision

1

Linguistics-111. Jupyter Notebook

1

awesome-llm-role-playing-with-persona. Awesome-llm-role-playing-with-persona: a curated list of resources for large language models for role-playing with assigned personas

1

LSGAN. LSGAN for image generation on anime dataset

1

AudioDec. An Open-source Streaming High-fidelity Neural Audio Codec

1

speech_quality. Jupyter Notebook

1

Codec-SUPERB. Audio Codec Speech processing Universal PERformance Benchmark

1

speech-language-model. A collection of papers related to speech language models

1
37
Apply