Classification of Web Pages as Evergreen Or Ephemeral Based on Content
Moonis Javed, Aly Akhtar, Akif Khan Yusufzai · 2015
Classification of web content is an interesting and widely pursued field of research in machine learning. Web classification could be done in various ways based upon the criteria chosen. Subjective classification involves classification of web pages based upon the subject to which these pages belong (say history, economics, politics, etc.). Another way of classifying web pages could be based upon the lifetime of these pages. A similar problem was introduced by Stumble Upon to classify the web pages as evergreen (larger lifetime) or ephemeral (not lasting for very long). The training dataset provided by Stumble Upon was a set of urls with some Meta information like the category of the page, html ratio, is news, along with some boilerplate code (like title and content) of the page. In this paper we have tried to use a novel methodology to classify these documents. The approach that we have used is a combination of text classification and other binary classification. Using this we have been able to get an overall accuracy of 88%.