Semantic similarity of short texts

Aminul Islam, Diana Zaiu Inkpen · Amsterdam studies in the theory and history of linguistic science. Series 4, Current issues in linguistic theory · 2009

This paper presents a method for measuring the semantic similarity of texts using a corpus based measure of semantic word similarity and a normalized and modified versions of the Longest Common Subsequence (LCS) string matching algorithm. Existing methods for computing text similarity have focused mainly on either large documents or individual words. In this paper, we focus on computing the similarity between two sentence or between two short paragraphs. The proposed method can be exploited in a variety of applications involving textual knowledge representation and knowledge discovery. Evaluation results on two different data sets show that our method outperforms several competing methods. Keywords Semantic similarity of words, similarity of short texts, corpusbased measures. 1.

Read the paper · More papers on PaperTik