Cuiabá, Mato Grosso - Brazil

Frederico S. Oliveira

Expert
@freds0

Researcher in the area of NLP, Ph.D. student at UFG, focusing on speech synthesis and recognition using deep learning and also professor at UFMT.

free-svc. [ICASSP 2025] FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion

95

data_augmentation_for_asr. A set of audio augmentation techniques to perform noise insertion in datasets used for Automatic Speech Recognition.

49

CML-TTS-Dataset. CML-TTS: A Multilingual Dataset for Speech Synthesis

36

katube. KATube is a tool to automate the process of creating datasets for training Text-To-Speech (TTS) and Speech-To-Text (STT) models. From a list of YouTube playlists or YouTube channels, KATube will generate dataset with audios and texts.

26

PTL-AI_Furnas_Dataset. PTL-AI Furnas Dataset: A Public Dataset for Fault Detection in Power Transmission Lines Using Aerial Images

24

fault_detection_power_transmission_lines. Tensorflow Object Detection API for fault detection at power transmission lines.

19

YourTTS. 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

16

kabooks. KABooks is a tool to automate the process of creating datasets for training Text-To-Speech (TTS) and Speech-To-Text (STT) models. Using audiobooks, KABooks will generate dataset with segmented audios and aligned texts.

13

BSpeech-MOS-Prediction. A model for predicting MOS that utilizes embeddings of supervised learning and self-supervised learning models, combined with embeddings of speaker verification models, to predict the MOS metric.

10

useful_audio_scripts. Some useful scripts for audio

9

capybara_dataset. This is a dataset composed of images of capybaras to be used for training a model for object detection

8

youtube_diarization. Download audios from youtube and execute diarization

6

Orpheus-TTS_unsloth_finetuning. Jupyter Notebook

5

speaker_clustering. Python

5

BRSpeech-Dataset. BRSpeech: A Portuguese Dataset for Speech Synthesis

5

CleanSpecNet. Python

4

CleanUNet2. Python

4

overlapping_voices_detector. Python

4

nvidia_tacotron2_multispeaker. Jupyter Notebook

4

tacotron2. Tacotron 2 - PyTorch implementation with faster-than-realtime inference adapted for brazilian portuguese.

4

CML-TTS-Toolkit. CML-TTS Conversion Tools

4

flask_fault_detection_power_transmission_lines. Frontend and backend separated object detection tf2 demo build with Flask, TensorFlow 2.x.

4

voice_gender_prediction. Python

3

hifi-gan. Python

2

wavegrad. A fast, high-quality neural vocoder.

2

algoritmos_e_estrutura_de_dados_material_didatico. Material Didático da disciplina Algoritmos e Estrutura de Dados

2

inteligencia_artificial_do_zero_ao_infinito. Material referente ao projeto sobre inteligência artificial com foco em visão computacional e detecão de objetos

2

yolov5. YOLOv5 🚀 in PyTorch > ONNX > CoreML > TFLite

1

XTTSv2-Finetuning-for-New-Languages. Python

1

MeloTTS. High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.

1

AI-Programming-using-Python. This repository contains implementation of different AI algorithms, based on the 4th edition of amazing AI Book, Artificial Intelligence A Modern Approach

1

hifi-gan2. Python

1

RVC-Demo. Python

1

MOSNet-pytorch. Python

1

Datasets-Portuguese-NLP. List of resources and tools developed with focus on Portuguese.

1

capybara_image_segmentation. This repository presents how to train your own Image Segmentator Using TensorFlow Object Detection API.

1
36
Apply