Data Structure Comparison Using Box Counting Analysis
Yukio Tominaga · Journal of Chemical Information and Computer Sciences · 1998
Box counting analysis was performed to visualize complex data structures of datasets. Two datasets were used as original datasets. One included 8000 samples and the other included 53 064 samples. Nine different selection methods were used to select subsets. The selection methods were as follows: maximum dissimilarity method, maximum similarity method, group averaging hierarchical clustering method, reciprocal nearest neighbor Ward hierarchical clustering method, k-mean nonhierarchical clustering methods with two types of seed points, Genetic algorithms (GAs), Kohonen networks, and cell-based method. The data structures of selected subsets were compared to that of the original dataset by using box counting analysis.