Feature Based Framework for Web Spam Filtering

Chirag Nathwani · SSRN Electronic Journal · 2020

Today widely used source of information is internet. Internet users use Search engines to find relevant information. Search engines worked based on page ranking algorithms. But, some web spam attempts to cheat this search algorithm by providing wrong metadata. We used decision tree for classifying whether particular page is spam or not. In this paper we compare results from 3 different data mining algorithms viz. Random Forest, J48 and LAD Tree. Experiments were carried out on standard data set WEBSPAM-UK2007. We added 5 new feature attributes in existing data set which were words found in spam pages and also implemented attribute reduction algorithm to remove redundant attributes from the data set. Reclassification using updated dataset improved precision by 2%. Moreover noticeable improvement of 20% was found in FP rate. Also attribute reduction helped in reducing build time by 8%.

Read the paper · More papers on PaperTik