Russian_subtitles_dataset. Preprocessing of the dataset of 347 subtitles for the TV series (thanks to Taiga Corpus) to build a word2vec model, JamSpell model, neural network training, chat bot training or in any other NLP task.

github.com/dbklim/Russian_subtitles_dataset

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.