Based on LSI and Dictionary Text Semantic Similarity Algorithm

WU Jun-hua · Coal Technology · 2010

The common problem in the fields of text clustering is that the conceptual similarity is ignored.We take thesaurus-based and corpus-based ontology to overcome this problem.A transformed latent semantic indexing(LSI) model which can appropriately capture the associated semantic similarity is proposed and demonstrated as corpus-based ontology.Experiments results show that the method apparently outperforms that with traditional similarity measures.

Read the paper · More papers on PaperTik