Fuzzy k-median Clustering Based on Hsim Function for the High Dimensional Data
Heng Zhao, Jimin Liang, Gaoyu Zhang · 2006
Hsim(x, y), a similarity measure function for high dimensional data is surveyed. The function can not only avoid the problem that Lk— norm leads to the non-contrasting behavior of distance in high dimensional space, but also adapt to both binary and numerical data. A fuzzy k-median clustering algorithm based on Hsim(x, y) is proposed. The algorithm uses Hsim(x, y) as the similarity measure of high dimensional data, and uses the approximated k-median algorithm optimize the center of cluster. The experiments indicate the algorithm is effective.