Improving Search Result Clustering by Enriching Snippets with Word2Vec Model

Nan Yang, Yaping Li, Qing Liu · 2017

Search Result Clustering (SRC) is an approach to solve the problems of web search engines under user's broad or ambiguous queries and no clues to find exact information in a long returned list. SRC groups the returned list and outputs a semantic structured organization to help users to find the desired information quickly. In a general way, the search engine results consist of concise information(snippet) about matching page documents. Because the snippet is short and carries few information, the clustering performance based on traditional TF-IDF weight is very low. An effective way to solve this problem is to enrich snippets according the semantic relationship from external text corpus. In this paper, we propose a new snippet enriching approach throughout word2vec model. The vector representations of words learned by word2vec models have been shown to carry semantic meanings and are useful in various Natural Language Processing(NLP) tasks. We propose to use the top-n similar words in word2vec model to enrich snippets and we still use traditional TF-IDF weights schema to select features. In order to demonstrate the effectiveness of our method, we design intensive experiments to evaluate new method and baseline methods. The result of the analysis shows that proposed method outperforms baseline approach significantly.

Read the paper · More papers on PaperTik