Binary codes for robust image hashing
Daniela Colţuc · 2013
In robust image hashing, the hash is obtained by binarizing a set of image features. Excepting some rare cases when the Gray code is used, the common solution for binarization is the natural binary code. By using as examples simulated and real data, we show that the choice of the binary code has effect on the hash properties and, consequently, on the collision probability. This probability is evaluated by estimating the mean of the hashes Hamming distance. For ideal hashes i.e., random and independent, the mean should be half of the hash length. Any correlation inside or between hashes has impact on the mean that falls down. We propose an algorithm for constructing codebooks with approximately constant Hamming distance between the consecutive codewords. With these codes, the mean bias reduces as the Hamming distance increases. The lowest bias is obtained for the maximum distance codebook.