PH-SSBM: Phrase Semantic Similarity Based Model for Document Clustering

Walaa Khaled Gad, Mohamed S. Kamel · 2009

In this paper, a novel document representation model the phrases semantic similarity based model (PHSSBM), is proposed. This model combines phrases analysis as well as words analysis with the use of WordNet as background knowledge to explore better ways of documents representation for clustering. The PH-SSBM assigns semantic weights to both document words and phrases. The new weights reflect the semantic relatedness between documents terms and capture the semantic information in the documents. The PH-SSBM finds similarity between documents based on matching terms (phrases and words) and their semantic weights. Experimental results show that the phrases semantic similarity based model (PH-SSBM) in conjunction with WordNet has a promising performance improvement for text clustering.

Read the paper · More papers on PaperTik