This is your work, valued

Baltimore, US

Helin Wang

Expert
@WangHelin1997

PhD at Johns Hopkins University

CapSpeech. CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

369

SoloSpeech. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

316

SSR-Speech. SSR-Speech: Towards Stable, Safe and Robust Zero-shot Speech Editing and Synthesis

155

SoloAudio. SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.

121

SpeechTasks. This is a list of speech tasks and datasets, which can provide training data for Generative AI, AIGC, AI model training, intelligent speech tool development, and speech applications.

83

MaskSpec. The Pytorch implementation of paper: Masked Spectrogram Prediction For Self-Supervised Audio Pre-Training

51

DuTa-VC. Source code and demo for INTERPSEECH 2023 paper: DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model

38

SpecAugment-plus. A Pytorch implementation of the paper : SpecAugment++: A Hidden Space Data Augmentation Method for Acoustic Scene Classification

34

Automatic_Speech_Annotator. Automatic speech annotator processing speech with voice activaty detection, overlapping speech detection, speaker diarization and automatic speech recognition

33

nnAudio2. Audio processing by using pytorch 1D convolution network (based on nnAudio). Gammatone Spectrogram and SpecAugmentation are now available on GPU.

21

DCASE-2020-Task1A-Code. A pytorch implementation of the paper : Acoustic Scene Classification with Multiple Decision Schemes.

20

AT-GCN. Pytorch implementation of the paper : Modeling Label Dependencies for Audio Tagging with Graph Convolutional Network

15

GL-AT. Pytorch implementation of the paper : A Global-local Attention Framework for Weakly Labelled Audio Tagging.

13

Aty-TTS. Aty-TTS: Improving fairness for spoken language understanding in atypical speech with Text-to-Speech

11

DCASE2021_Task6_PKU. This is the code of PKU team for DCASE 2021 Task 6.

9

Speech-Captioning-Dataset. Python

8

FPNet. A signal segmentation method of CNN for audio event classification

7

CNN-model-and-visualization. A CNN model (RseNet) for image classification( CIFAR-10), including filter and output of layers visualization.

6

project2021. PKU team for 2021 project 'Guangchangwu detection'.

3

Du-N2DVC-Demo. HTML

3

CapSpeech-demo. JavaScript

3

SCNN. SincConv layer using in AED and ASC

3

DCASE2020-Task6-PKU. A Pytorch implementation of the DCASE2020 Task6 by PKU team : Automated Audio Captioning With Temporal Attention

2

CommonVoice. Python

2

Babycry-sound-detection. PyTorch implementations of neural network models for Babycry sound detection, including training process and test demo. Based on DCASE2017 Task2: Detection of rare sound events.

2

VGGSound-Mix. Python

2

dcase2019_1D. Dcase2019 Task1a using audio feature module.

1

helinwang. JavaScript

1

SoloSpeech-Demo. JavaScript

1

ATReSN-Net. Capturing attentive temporal relations in semantic neighborhood for ASC

1

Attention-based_Atrous_CNN. Pytorch code for the paper 'Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic Scenes', by Zhao Ren, Qiuqiang Kong, Jing Han, Mark Plumbley, Björn Schuller.

1

SSR-Speech-Demo. https://wanghelin1997.github.io/SSR-Speech-Demo/

1

Pytorch-audio_feature. Audio feature extraction in Pytorch module.

1

Your-Stable-Audio. Stable Audio UnOffical Implementation: Latent Diffusion for Audio Generation

1

WaveMsNet. Python

1

ESResNet. Source code for models described in the paper "ESResNet: Environmental Sound Classification Based on Visual Domain Models" (https://arxiv.org/abs/2004.07301)

1

lgtfb-en. Learnable Gammatone Filterbank (LGTFB) and Equal-loudness Normalization (EN)

1