Document clustering based on web search hit counts
Masaya Kaneko, Shusuke Okamoto, Masaki Kohana, You Inayoshi · International Journal of Business Intelligence and Data Mining · 2013
This paper describes a web mining method for clustering research documents automatically. Web hit counts of AND-search for two words are used to form a document feature vector. Target documents are clustered using the k-means clustering method twice, in which cosine similarity is used to calculate the distance measure.