Detecting translations of the same text and data with common source

Kostadin Koroutchev, Manuel Cebrián · Journal of Statistical Mechanics Theory and Experiment · 2006

Compression based similarity distances have the main drawback of needing the same coding scheme for the objects to be compared. In some situations, there exists significant similarity with no literal shared information: text translations, different coding schemes, etc. To overcome this problem, we present a similarity measure that compares the redundancy structure of the data extracted by means of a Lempel–Ziv compression scheme. Each text is represented as a graph in which vertices are text positions and edges represent shared information; with our measure, two texts are similar if they have the same referential topology when compressed. In this paper we give empirical evidence and a phenomenological explanation that this new measure is a robust indicator, detecting similarity between data coded in different languages. We also regard a textual data without any structure, but with a common source, and find that we can detect such data and distinguish this situation from the previous one.

Read the paper · More papers on PaperTik