Automated building and analysis of Ukrainian Twitter corpus for toxic text detection

Kateryna Bobrovnyk · Electronic Scientific Archive (Lviv Polytechnic) · 2019

Toxic text detection is an emerging area of study in Inter-net linguistics and corpus linguistics. The relevance of the topic can be explained by the lack of Ukrainian social media text corpora that are publicly available. Research involves building of the Ukrainian Twitter corpus by means of scraping; collective annotation of 'toxic/non-toxic' texts; construction of the obscene words dictionary for future feature engineering; and models training for the task of text classi cation (com-paring Logistic Regression, Support Vector Machine, and Deep Neural Network).

Read the paper · More papers on PaperTik