Clustering Interval-valued Data Using Principal Components
Jiankun Zhu, Lynne Billard · Journal of Statistical Theory and Practice · 2025
Abstract Interval-valued observations are examples of symbolic data. This article adapts the concepts of principal component analysis and clustering algorithms to develop a new methodology for interval-valued data. Chavent’s [10, 11] monothetic divisive clustering algorithm has been used extensively for clustering interval-valued observations. To overcome some limitations of this algorithm, three new algorithms are proposed herein, one using the Chavent center based ordering idea but applied to principal components of each hypercube, one as a double ordering criteria using both interval endpoints, and a third as a mixed-strategy double algorithm that is based on the principal components criteria applied to both interval endpoints. Simulations show the proposed algorithms outperform previous methods; real data are analysed.