Lexical Chains and Sliding Locality Windows in Content-based Text Similarity Detection
Thade Nahnsen, Özlem Uzuner, Boris Katz · 2005
We present a system to determine content similarity of documents. More specifi-cally, our goal is to identify book chapters that are translations of the same original chapter; this task requires identification of not only the different topics in the docu-ments but also the particular flow of these topics. We experiment with different rep-resentations employing n-grams of lexical chains and test these representations on a corpus of approximately 1000 chapters gathered from books with multiple paral-lel translations. Our representations in-clude the cosine similarity of attribute vectors of n-grams of lexical chains, the cosine similarity of tf*idf-weighted key-words, and the cosine similarity of un-weighted lexical chains (unigrams of lexical chains) as well as multiplicative combinations of the similarity measures produced by these approaches. Our results identify fourgrams of unordered lexical chains as a particularly useful representa-tion for text similarity evaluation. 1