Evaluation of text classification techniques for inappropriate web content blocking
Igor Vitalievich Kotenko, Andrey Alexeevich Chechulin, Dmitry Komashinsky · 2015
The paper is devoted to the issues of automated categorization of textual information which can be applied in the systems intended to block inappropriate content. The approach used for feature selection and construction is proposed. The text mining methods used for research (Decision Tree classifiers) are analyzed. Besides that, the techniques of Web sites analysis that provide information in different languages are suggested. The aspects of collection and analysis of text features required for classification in certain categories are investigated. Results of experiments on analysis of text correspondence to different categories are given. The classification quality is evaluated. The text classification component, developed as a result of this paper, is intended for realization in F-Secure systems aiming to block inappropriate web content.