Edge-based categorical clustering for data with related features
Weronika Łazarz, Agnieszka Nowak-Brzezińska · Procedia Computer Science · 2024
Supervised methods have been the leading methods for inferring from data, but unsupervised models, particularly with data clustering, are becoming more common in business processes. Cluster analysis methods are extremely effective when we do not have a clearly defined decision-making class, or the data contain anomalies that are difficult to identify. A particularly big problem arises with the efficient processing of categorical data. Such data type is, for example, amino acid sequences, where their co-occurrence is extremely important. Therefore, we propose a new approach to clustering categorical sequential data based on graph edges -GECC. This algorithm not only identifies clusters well but also, due to its implementation based on a graph structure, it is fast and does not burden memory resources. This article will show that our method perfectly reflects the natural data clusters.