Detection of near dublicates in tables based on the locality-sensitive hashing method and the nearest neighbor method
Петро Лізунов, Andrii O. Biloshchytskyi, Alexander Kuchansky, Svitlana Biloshchytska, Larysa Chala · Eastern-European Journal of Enterprise Technologies · 2016
to be one of the steps of the process of their mapping and visualization.In general, a table consists of orderly arranged rows and columns.The term "row" can be defined also as vector, record or tuple.The usual term "column" is often interpreted as field, attribute or property that has a specified title or a name.This name may a priori consist of a word or a sequence of words, be presented by numeric values, a formula (formulas), or a date.The intersection of the particular row and column defines a cell, which is filled with content.In general, the presentation of tables is very diverse: elements of the tables can be grouped according to certain attributes, split into segments, nested with varying degrees of nesting, contain annotations, objects of the type of formula, date, etc.In the case of hiding plagiarism, the original table can be relatively easily modified: lines and columns transposed, 4