Ranking the blocking keys for data de-duplication in information systems

Asif Sohail, Syed Waqar Jaffry · International Journal of Business Information Systems · 2025

Data de-duplication is an essential activity in data integration and data cleansing. It identifies and removes the disguised duplicates in a dataset. Blocking is an established technique for reducing the inherent quadratic complexity of de-duplication. Blocking gathers the potential matching records in the same block on the basis of a blocking key. The results of blocking fluctuate considerably when different blocking keys are employed. Hence, it becomes extremely important to select an appropriate blocking key for maximising the efficacy and efficiency of blocking. The proposed technique ranks the attributes of a dataset with respect to their usability as a blocking key. We have introduced a novel correlation measure called R-score for computing correlation between gold rankings and computed rankings of the blocking keys. The proposed technique is evaluated using benchmark datasets and the experimental results confirm that the proposed technique outperforms the existing techniques.

Read the paper · More papers on PaperTik