Detecting Fake Sites based on HTML Structure Analysis
Jiachang Xu, Kilho Shin, Yulu Liu · 2016
Fake sites are serious threats for both consumers and merchants. They mimic authentic sites and attempt to obtain money by fraud, steal personal information such as credit card numbers from consumers, ruin merchants' reputation and so on. Usually, it is difficult for consumers to distinguish between authentic sites and fake sites. In this paper, we propose a method to detect fake sites with high accuracy based on analyzing structures of HTML source codes of websites. For the analysis, we leverage a combination of a tree kernel and the support vector machine (SVM). We first test 32 different types of tree kernels with small datasets, and then, select a few kernels that have exhibited the best accuracy scores. Further, we test them with larger datasets, and finally, focus on a single kernel that has shown the best accuracy and speed. The AUC, F-measure, Accuracy scores of the kernel exceed 0.97 with a large dataset with 3001 instances, and we conclude that the method can perform at the practical level.