Efficient Algorithm for Near Duplicate Documents Detection
Gaudence Uwamahoro, Zuping Zhang · 2013
Identification of duplicates or near duplicate documents in a set of documents is one of the major problems in information re trieval. Several methods to detect those documents have been proposed but their relevance is still an issue. In this paper we propose an algorithm based on word position which provides a reduced candidate size to search in and increases efficiency and effectiveness for partial documents relevance. In our experiments the results show that during search process for the query the candidate size has reduced up to 12% of the size of set of documents which leads to a decreased time in searching. The results also have shown a higher accuracy thus helping help the user not to waste time on waiting for a query and getting unwanted documents.