Mining for high complexity regions using entropy and box counting dimension quad-trees
Rosanne Vetro, Wei Ding, Dan A. Simovici · 2010
This paper introduces an algorithm for capturing high complexity regions of a data domain. In this work, we focus on domains in R2. In particular, we analyze 2-dimensional image domains. Two different methods for mining are considered. The first method performs an information-theoretic analysis based on entropy to find diverse areas. The second method applies the concept of box-counting dimension related to fractal geometry. We propose the use of a quad-tree as main search structure where complex areas are represented by leaves with high feature values at the highest level on the tree. Nodes that refer to specific sub-domains are split when the level of the analyzed feature exceeds a chosen threshold. The relationship between the threshold and the number of pixels located in high value feature sub-domains at the highest level on the resultant quad-tree is demonstrated on test images for both methods. Experimental results also show the relation between the former measurements and characteristics of the images. Finally, we identify a correlation between the methods presented.