A Novel Technique Using Multiple K-Shingling Based Weighted Dissimilarity Score for Web Content Outlier Mining

Liyakath Khan, Mohammed Ahmed, Husni Almistarihi · International journal of intelligent engineering and systems · 2019

The technological evolution of Internet and Web along with several applications leads to the problem of redundancy as the documents are unanimously forwarded and are stored in several servers and platforms.Recently, not only duplicate documents but also near-duplicate documents affect the performance of the search results.The main objective of this paper is to provide significant documents by eliminating the redundancy and near redundancy documents present in the web search results.The proposed model comprises of two phases such as pre-processing phase and dissimilarity computation phase.For dissimilarity computation, the proposed model employs multiple kshingling based dissimilarity score to identify the duplicate and near-duplicate documents which are considered as the outliers present in the set of input web documents.The proposed model has been evaluated using several experimental analysis.As there are no real datasets available for duplicate detection, datasets have been created and the performance evaluation is carried out with the created datasets.Several statistical analysis has been made wherein the average specificity, sensitivity, precision, and accuracy are 87%, 93%, 80%, and 92% respectively.The comparative analysis has also been made with various existing methods, in which the proposed model provides better results than existing methods in removing near-duplicates.The proposed multiple k-shingling based weighted dissimilarity model effectively detects the duplicates and near-duplicates when the number of outliers is minimum.

Read the paper · More papers on PaperTik