A Pure URL-Based Genre Classification of Web Pages

Chaker Jebari · 2014

In this paper, we propose a new approach for multi-label genre classification of web pages that exploits character n-grams extracted from the URL of the web page rather than its content. Using only the URL reduces the time needed for feature extraction since it does not need to download the content of the web page. Our approach deals with the complexity of web pages because it uses a multi-label classification where each web page can be assigned to more than one genre. Moreover, our approach implements a new weighting technique that exploits the structure of the URL. Experiments conducted on a known multi-label dataset show that our approach achieves encouraging results.

Read the paper · More papers on PaperTik