Document clustering based on web search hit counts

Masaya Kaneko, Shusuke Okamoto, Masaki Kohana, You Inayoshi · International Journal of Business Intelligence and Data Mining · 2013

This paper describes a web mining method for clustering research documents automatically. Web hit counts of AND-search for two words are used to form a document feature vector. Target documents are clustered using the k-means clustering method twice, in which cosine similarity is used to calculate the distance measure.

Read the paper · More papers on PaperTik