K-mixed prototypes
Rahmah Brnawy, Nematollaah Shiri · 2019
Many real-life applications involve data with mixed numeric and categorical values. While the notion of similarity/distance measure is well defined for numeric values, defining the distance between categorical values is not as straightforward, mainly because categorical values have no order. The situation is even less clear for clustering data with mixed numeric and categorical values. Among the categorical types of attributes, nominal versus binary types can further be distinguished. Also, for a categorical binary attribute, the contributions of the two types of values may not be of equal importance. Whereas existing similarity measures for categorical values do not recognize these differences, we show that recognizing them can yield improved clustering results. To do this, we first extend an existing clustering method proposed originally for pure categorical data and then we develop k-mixed prototypes for clustering data with mixed numeric and categorical values. Our technique requires only the number of clusters as the input parameter, and it recognizes the differences mentioned. The results of our numerous experiments, using real-life benchmark data, indicate significant increase in accuracy and efficiency of the proposed k-mixed prototype, compared to existing clustering methods.