On Preprocessing the Data for Improving Sexual Predator Detection : Anonymous for review

Parisa Rezaee Borj, Kiran Bylappa Raja, Patrick Bours · 2020

Sexual predator detection and predatory message identification are critical to avoid under-aged children from being abused online. In this paper, we investigate different feature extraction approaches for predatory detection. While the previous results indicate good accuracy on predatory conversation detection, there is a missing investigation on the robustness of feature space. Further, we also show the impact of preprocessing on data to improve the performance of predator identification and predatory message classification. Various types of the bag of words features, including binary, term frequency, and TF-IDF representation are investigated on the publicly available PAN 2012 competition dataset for predator identification. Further, to cover the relationship between the words in the text analysis, the GloVe feature set is also investigated for word embedding features. With the set of preprocessing of data, we illustrate the improvement in detecting predatory conversation with an accuracy of 0.994 and F1-score of 0.964.

Read the paper · More papers on PaperTik