Ph.D.ing | Research Intern @QwenLM | Prev: @ByteDance-Seed @alibaba-damo-academy @microsoft @megvii-research
emotion2vec. [ACL 2024] Official PyTorch code for extracting features and training downstream models with emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
1.2kSpeech-Resources. 语音方向实验室/公司/资源/实习等,欢迎推荐或自荐
609MMAR. [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
214Awesome-Speech-Pretraining. Paper, Code and Statistics for Self-Supervised Learning and Pre-Training on Speech.
212Awesome-Speech-Language-Model. Paper, Code and Resources for Speech Language Model and End2End Speech Dialogue System.
202Omni-Captioner. [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.
142MMAE. MMAE: A Massive Multitask Audio Editing Benchmark
102MT4SSL. [INTERSPEECH 2023 Best Paper Shortlist] Official implementation for MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets
45Awesome-Speech-Generation. Paper, Code and Statistics for Speech Generatation.
10HDRR. Official implementation for Hierarchical Deep Residual Reasoning for Temporal Moment Localization
9Multimodal_Visualization_Framework. 一款时域语言定位可视化框架
6pre-train-dockerfile. An Intro to set up your Speech Docker environment and debug using VSCode
4amlt. A repo for amlt examples.
3Industrial_robot. 使用MOTOMAN-GP25完成简易物流平台的搭建,通过视觉模块—通讯模块—控制模块联动完成物流分拣功能。
2DL-NLP-Readings. My Reading Lists of Deep Learning and Natural Language Processing
1espnet. End-to-End Speech Processing Toolkit
1MovieChat. 🎬💭 chat with over 10K frames of video!
1