Research on Spam Web Page Detection Based on Unbalanced Data Processing

Xiaxia Yang, Xuxia Huang, Yanjun Wang · 2021

Most Internet users are used to paying attention to a few websites with the top ranking of search engines. According to statistics, 95% of users are only interested in the search results of the first five pages, and a large number of websites with the lower ranking are selectively ignored by users. The purpose of this paper is to make use of the discrimination results of multiple classifiers, and build an effective combination classifier for spam web page recognition from several aspects such as feature selection, feature segmentation and classifier difference measurement. The new two-view features are combined in different ways to generate single-view data, and this set of data is used as training data to construct classification algorithm. Experimental results show that the proposed method can effectively identify illegal publishers who commit fraudulent clicks, and meet the requirements of click fraud detection in online advertising.

Read the paper · More papers on PaperTik