Near duplicate text detection using graph depiction

Marios Poulos · 2016

In this paper, a new identification technique based on a Kendal Rank Correlation Model is described. The method is based on the exploitation of a text vector model in the area of the text similarity detection issue. This study is divided into two stages: In the first stage the Text Vector Model (TVM) is described. In the second stage the Kendal rank correlation coefficients, which are obtained from the processing text vector model, are fitted on a 4th order polyonym using least square technique. Then the graph of the third degree polynomial consists of the criterion of the plagiarism detection. Also, this method is based on transformation of cosine similarity process and is compared with the classical techniques Jaccard and cosine similarities. The experimental results showed that this method is very sensitive in the non-similar small texts in contrast to the similar. This observation should be valuable to detect similar small texts which have been submitted in the paraphrase and quote or other similar text processing.

Read the paper · More papers on PaperTik