Online clustering with interpretable drift adaptation to mixed features

Flavio Corradini, Vincenzo Nucci, Marco Piangerelli, Barbara Re · Intelligent Systems with Applications · 2025

In the era of big data, the rapid pace and variability of information have become increasingly evident, particularly in areas like seasonal trends and manufacturing processes. The dynamic nature of the environments that produce these data means that their behavior is time-dependent. Consequently, treating data streams as static entities is no longer effective. This has led to the concept of data drift, which refers to shifts in data distribution over time. Stream processing algorithms are designed to detect these changes promptly and adjust to the newly emerging data patterns. In our research, we introduce FURAKI, an innovative online clustering algorithm that incorporates drift detection. It employs a binary tree structure and is capable of handling both single-feature and mixed-feature data from unbounded streams. We conducted extensive testing of FURAKI against state-of-the-art algorithms using various datasets. Our findings reveal that FURAKI outperforms the state-of-the-art algorithms in the considered datasets. • FURAKI: Unsupervised online clustering for mixed data with concept drift handling. • Drift Types: Detects abrupt, recurrent, incremental, and gradual concept drifts using KDE and G-tests. • Interpretable Model: Uses binary tree to show features driving changes in clustering. • Performance: Beats SOTA in F1-score and ARI on synthetic and real-world data. • Mixed Features: Handles numeric and categorical features without transformations.

Read the paper · More papers on PaperTik