This is your work, valued

Ziyang Ma

Elite
@ddlBoJack

Ph.D.ing | Research Intern @QwenLM | Prev: @ByteDance-Seed @alibaba-damo-academy @microsoft @megvii-research

emotion2vec. [ACL 2024] Official PyTorch code for extracting features and training downstream models with emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

1.2k

Speech-Resources. 语音方向实验室/公司/资源/实习等,欢迎推荐或自荐

609

MMAR. [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

214

Awesome-Speech-Pretraining. Paper, Code and Statistics for Self-Supervised Learning and Pre-Training on Speech.

212

Awesome-Speech-Language-Model. Paper, Code and Resources for Speech Language Model and End2End Speech Dialogue System.

202

Omni-Captioner. [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.

142

MMAE. MMAE: A Massive Multitask Audio Editing Benchmark

102

MT4SSL. [INTERSPEECH 2023 Best Paper Shortlist] Official implementation for MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets

45

Awesome-Speech-Generation. Paper, Code and Statistics for Speech Generatation.

10

HDRR. Official implementation for Hierarchical Deep Residual Reasoning for Temporal Moment Localization

9

Multimodal_Visualization_Framework. 一款时域语言定位可视化框架

6

pre-train-dockerfile. An Intro to set up your Speech Docker environment and debug using VSCode

4

amlt. A repo for amlt examples.

3

Industrial_robot. 使用MOTOMAN-GP25完成简易物流平台的搭建,通过视觉模块—通讯模块—控制模块联动完成物流分拣功能。

2

DL-NLP-Readings. My Reading Lists of Deep Learning and Natural Language Processing

1

espnet. End-to-End Speech Processing Toolkit

1

MovieChat. 🎬💭 chat with over 10K frames of video!

1