YQ

Elite
@ranchlai

mandarin-tts. Chinese Mandarin tts text-to-speech 中文 (普通话) 语音 合成 , by fastspeech 2 , implemented in pytorch, using waveglow as vocoder, with biaobei and aishell3 datasets

477

speaker-verification. Speaker verification using ResnetSE (EER=0.0093) and ECAPA-TDNN

97

awesome-speaker-embedding. A curated list of speaker-embedding speaker-verification, speaker-identification resources.

52

lectures. B站视频课程配套资料

40

pinyin2hanzi. 拼音转汉字, convert pinyin to 汉字 using deep networks

23

nlpcc2023-shared-task-diaASQ. NLPCC2023 shared-task DiaASQ first-place solution. (NLPCC2023对话式细粒度情感识别大赛第一名方案)

15

adversarial_example. transform sky to cat

7

clip.paddle. OpenAI clip implementation in PaddlePaddle

5

sound_classification. pretrained sound classification models on esc-50. acc=0.937

5

wav2vec-2.0. Wav2vec2 English speech recognition in PaddlePaddle

4

gsm8k. gsm8k in json format

3

quantizations. A collection of quantization recipes for various large models including Llama-2-70B, QWen-14B, Baichuan-2-13B, and more.

3

VocGAN. VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network

2

bitsandbytes. 8-bit CUDA functions for PyTorch

2

awesome-sound. Awesome acoustic(sound) scenes and events classification and detection, papers, code, toolkits, frameworks and libraries

2

speaker-benchmarks. Testing SOTA open-source Voxceleb speaker models on Aishell-3

2

ColossalAI. Making big AI models cheaper, easier, and scalable

1

GODEL. Large-scale pretrained models for goal-directed dialog

1

DFRF. [ECCV2022] The implementation for "Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis".

1

Awesome-generative-recommendation-using-llm. This repository contains a collection of papers, blog posts, and tutorials on generative recommendation using large language models

1

DeepSearch. Minimum code to implement Deep search(深度搜索) using deepseek-V3 and Jina API

1

qlora. QLoRA: Efficient Finetuning of Quantized LLMs

1

peft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

1

dcase2021_task1b. Python

1

trlx. A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)

1

Megatron-DeepSpeed. Ongoing research training transformer language models at scale, including: BERT & GPT-2

1

insurance-demo. Building a bot to handle general tasks for insurance.

1

dstc11-track5. DSTC11 Track 5 - Task-oriented Conversational Modeling with Subjective Knowledge

1

ConvLab-3. Python

1

GPT2-chitchat. GPT2 for Chinese chitchat/用于中文闲聊的GPT2模型(实现了DialoGPT的MMI思想)

1

speechbrain. A PyTorch-based Speech Toolkit

1

waveglow. A Flow-based Generative Network for Speech Synthesis

1

wenet. Production First and Production Ready End-to-End Speech Recognition Toolkit

1

fairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.

1

pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration

1

preprocessing. Audio preprocessing using PaddleAudio

1

AudioCLIP. Source code for models described in the paper "AudioCLIP: Extending CLIP to Image, Text and Audio" (https://arxiv.org/abs/2106.13043)

1

espnet. End-to-End Speech Processing Toolkit

1

rasa. 💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants

1

audio. Data manipulation and transformation for audio signal processing, powered by PyTorch

1

models. Pre-trained and Reproduced Deep Learning Models (『飞桨』官方模型库,包含多种学术前沿和工业场景验证的深度学习模型)

1

ConvLab-2. ConvLab-2: An Open-Source Toolkit for Building, Evaluating, and Diagnosing Dialogue Systems

1
42
Apply