Using main content extraction to improve performance of Vietnamese web page classification

Nguyen Minh Trung, Nguyen Duc Tam, Nguyễn Hồng Phương · 2011

Web page classification is the process of categorizing a web page into one or more classes which have been predetermined. If we remove all HTML tags from a web page, then this process can be considered as a text classification problem. However, this approach does not achieve high precision due to noisy contents, which always exist in regular HTML documents. To address this problem, we propose using a content extraction method to extract the main contents of the web pages and use them for the classification task. Experimental results show that the proposed method significantly improves the precision of the Vietnamese web page classification from 71% to 80%. It also indicates that context features such as the anchor texts of reference links and the contents of tags "TITLE" can use as a good summarization for web page contents.

Read the paper · More papers on PaperTik