The Identification of Mass Media by Text Based on the Analysis of Vocabulary Peculiarities Using Support Vector Machines

Maksym Lupei, Олександр Володимирович Міца, Vasyl Sharkan, Sabolch Vargha, Василь Михайлович Горбачук · 2022 International Conference on Smart Information Systems and Technologies (SIST) · 2022

The study proposes the approach to identifying whether a text belongs to certain media with the help of the support vector machine. Based on the results of studying the texts of the Ukrainian media “Ukrainska Pravda” and “Puliteka”, it was determined that, depending on a change in parameters or machine learning architecture, the result improves or worsens. It has been established that on the basis of a database containing 6000 text fragments, evenly distributed in alphabetical order, the reliability coefficient of identification of texts as belonging to the specific mass media using a SVM (SVC / SVR) is about 0.99. It has been determined that at the lexical language level the text is defined as belonging to certain media, mainly due to lexico-semantic and, to a lesser extent, lexico-stylistic peculiarities.

Read the paper · More papers on PaperTik