Ambiguity-Aware Document Similarity

Fabrizio Caruso, Giovanni Giuffrida, Diego Reforgiato, Giuseppe Tribulato, Calogero G. Zarba · International Journal on Natural Language Computing · 2013

In recent years, great advances have been made in the speed, accuracy, and coverage of automatic word sense disambiguator systems that, given a word appearing in a certain context, can identify the sense of that word.In this paper we consider the problem of deciding whether same words contained in different documents are related to the same meaning or are homonyms.Our goal is to improve the estimate of the similarity of documents in which some words may be used with different meanings.We present three new strategies for solving this problem, which are used to filter out homonyms from the similarity computation.Two of them are intrinsically non-semantic, whereas the other one has a semantic flavor and can also be applied to word sense disambiguation.The three strategies have been embedded in an article document recommendation system that one of the most important Italian ad-serving companies offers to its customers.

Read the paper · More papers on PaperTik