An Information-Theoretic Sentence Similarity Metric

John G. Mersch, Rongling Lang · The Florida AI Research Society · 2015

We describe an information theoretic-based metric for sentence similarity. The method uses the information content (IC) of dependency triples using corpus statistics generated by processing the Open American National Corpus (OANC) with the Stanford Parser. We define the similarity of two sentences as a function of (1) the similarity of their constituent dependency triples, and (2) the position of the triples in their respective dependency trees. We compare results of the algorithm to human judgments of similarity of 1725 sentence pairs.

Read the paper · More papers on PaperTik