Simplistic Hashing for Building a Better Bloom Filter on Randomized Data
Ahmad Ali Iqbal, Maximilian Ott, Aruna Prasad Seneviratne · 2010
User demands to have access to complete and accurate information requires integration of data from distributed stores. Those stores provide dynamically changing data that could be partially redundant because of many intentional or unintentional reasons. The unintentional reasons could be the way the data was collected by those information stores and intentional reasons could be replication or the nature of the content description language. Whatever the reason is, a need for a filtering mechanism during information retrieval so that redundancy of data could be removed before transmitting on the network arises. This paper proposes an improvement to the randomized redundant data filtering by the support of an efficient hashing algorithm. We evaluate these hashing algorithms for building a Bloom filter on randomized data.