Web Spam Challenge 2007 Track II Secure Computing Corporation Research

Yuchun Tang, Yuanchen He, Sven Krasser, Paul Judge · 2007

Abstract. To discriminate spam Web hosts/pages from normal ones, text-based and link-based data are provided for Web Spam Challenge Track II. Given a small part of labeled nodes (about 10%) in a Web linkage graph, the challenge is to predict other nodes ’ class to be spam or normal. We extract features from link-based data, and then combine them with text-based features. After feature scaling, Support Vector Machines (SVM) and Random Forests (RF) are modeled in the extremely high dimensional space with about 5 million features. Stratified 3-fold cross validation for SVM and out-of-bag estimation for RF are used to tune the modeling parameters and estimate the generalization capability. On the small corpus for Web host classification, the best F-Measure value is 75.46 % and the best AUC value is 95.11%. On the large corpus for Web page classification, the best F-Measure value is 90.20 % and the best AUC value is 98.92%.

Read the paper · More papers on PaperTik