An Information Classification Approach Based on Knowledge Network

Huakang Li, Guozi Sun, Bei Xu, Li Li, Jie Huang, Keita Tanno, Wenxu Wu, Changen Xu · 2014

Numerous critical Internet applications with high-quality services, such as Web directory, search engine, Web crawler, recommendation system and user profile detector, etc. Almost depend on the efficient and accurate of web page classification system. Traditional supervised or semi-supervised machine learning methods become more and more difficult to adapt to the explosive Internet information. This paper proposed a web page classification method based on the topological structure of Wikipedia knowledge network. The kinship-relation association based on content similarity was proposed to solve the unbalance problem when a category node inherited the probability from multiple fathers. We used N-gram based on Wikipedia words to extract the keywords from web page, and introduce Bayes classifier to estimate the page class probability. Experimental results shown that the proposed method has very good scalability, robustness and reliability for different web pages.

Read the paper · More papers on PaperTik