Discovering Similarity Across Heterogeneous Features

Vandana Pursnani Janeja, Josephine Namayanja, Yelena Yesha, Anuja Kench, Vasundhara Misal · International Journal of Data Warehousing and Mining · 2020

The analysis of both continuous and categorical attributes generating a heterogeneous mix of attributes poses challenges in data clustering. Traditional clustering techniques like k-means clustering work well when applied to small homogeneous datasets. However, as the data size becomes large, it becomes increasingly difficult to find meaningful and well-formed clusters. In this paper, the authors propose an approach that utilizes a combined similarity function, which looks at similarity across numeric and categorical features and employs this function in a clustering algorithm to identify similarity between data objects. The findings indicate that the proposed approach handles heterogeneous data better by forming well-separated clusters.

Read the paper · More papers on PaperTik