Document classification based on web search hit counts
Masaya Kaneko, Shusuke Okamoto, Masaki Kohana, You Inayoshi · 2012
This paper describes a web mining method to classify research documents automatically. Web hit counts of AND-search on two words are used to form a document vector. Target documents are classified with a result of k-means clustering method, in which cosine similarity is used to calculate a distance.