This is your work, valued
mandarin-tts. Chinese Mandarin tts text-to-speech 中文 (普通话) 语音 合成 , by fastspeech 2 , implemented in pytorch, using waveglow as vocoder, with biaobei and aishell3 datasets
477speaker-verification. Speaker verification using ResnetSE (EER=0.0093) and ECAPA-TDNN
97awesome-speaker-embedding. A curated list of speaker-embedding speaker-verification, speaker-identification resources.
52lectures. B站视频课程配套资料
40pinyin2hanzi. 拼音转汉字, convert pinyin to 汉字 using deep networks
23nlpcc2023-shared-task-diaASQ. NLPCC2023 shared-task DiaASQ first-place solution. (NLPCC2023对话式细粒度情感识别大赛第一名方案)
15adversarial_example. transform sky to cat
7clip.paddle. OpenAI clip implementation in PaddlePaddle
5sound_classification. pretrained sound classification models on esc-50. acc=0.937
5wav2vec-2.0. Wav2vec2 English speech recognition in PaddlePaddle
4gsm8k. gsm8k in json format
3quantizations. A collection of quantization recipes for various large models including Llama-2-70B, QWen-14B, Baichuan-2-13B, and more.
3VocGAN. VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
2bitsandbytes. 8-bit CUDA functions for PyTorch
2awesome-sound. Awesome acoustic(sound) scenes and events classification and detection, papers, code, toolkits, frameworks and libraries
2speaker-benchmarks. Testing SOTA open-source Voxceleb speaker models on Aishell-3
2ColossalAI. Making big AI models cheaper, easier, and scalable
1GODEL. Large-scale pretrained models for goal-directed dialog
1DFRF. [ECCV2022] The implementation for "Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis".
1Awesome-generative-recommendation-using-llm. This repository contains a collection of papers, blog posts, and tutorials on generative recommendation using large language models
1DeepSearch. Minimum code to implement Deep search(深度搜索) using deepseek-V3 and Jina API
1qlora. QLoRA: Efficient Finetuning of Quantized LLMs
1peft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
1dcase2021_task1b. Python
1trlx. A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)
1Megatron-DeepSpeed. Ongoing research training transformer language models at scale, including: BERT & GPT-2
1insurance-demo. Building a bot to handle general tasks for insurance.
1dstc11-track5. DSTC11 Track 5 - Task-oriented Conversational Modeling with Subjective Knowledge
1ConvLab-3. Python
1GPT2-chitchat. GPT2 for Chinese chitchat/用于中文闲聊的GPT2模型(实现了DialoGPT的MMI思想)
1speechbrain. A PyTorch-based Speech Toolkit
1waveglow. A Flow-based Generative Network for Speech Synthesis
1wenet. Production First and Production Ready End-to-End Speech Recognition Toolkit
1fairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
1pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
1preprocessing. Audio preprocessing using PaddleAudio
1AudioCLIP. Source code for models described in the paper "AudioCLIP: Extending CLIP to Image, Text and Audio" (https://arxiv.org/abs/2106.13043)
1espnet. End-to-End Speech Processing Toolkit
1rasa. 💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
1audio. Data manipulation and transformation for audio signal processing, powered by PyTorch
1models. Pre-trained and Reproduced Deep Learning Models (『飞桨』官方模型库,包含多种学术前沿和工业场景验证的深度学习模型)
1ConvLab-2. ConvLab-2: An Open-Source Toolkit for Building, Evaluating, and Diagnosing Dialogue Systems
1