Clustering methods for categorical time series and sequences : a scoping review
O. Khalifa, Alan Balendran, Viet-Thi Tran, F. Petit · Journal of Epidemiology and Population Health · 2026
OBJECTIVE: To provide an overview of clustering methods for categorical time series (CTS), a data structure common in epidemiology, sociology, biology, and marketing, and to support method selection according to data characteristics. MATERIALS AND METHODS: We searched PubMed (via MEDLINE), Web of Science, and Google Scholar up to November 2024 for articles proposing and evaluating CTS clustering techniques. Methods were classified into three families-distance-based, feature-based, and model-based-and assessed for their ability to address challenges such as variable sequence length, multivariate data, continuous time, missing data, covariates, and large data volumes. RESULTS: Of 14,607 records retrieved, 124 articles describing 129 methods were included. Distance-based approaches, especially those using Optimal Matching, were most common, with 56 methods. We found 28 model-based methods, which covered a broader range of complex data structures such as multivariate data, continuous time and time-invariant covariates. We recorded 45 feature-based approaches, which were on average more scalable but less flexible. Fewer than half of the methods provided public implementations. A searchable Web application ( https://cts-clustering-scoping-review-7sxqj3sameqvmwkvnzfynz.streamlit.app/ ) was developed to support method selection. DISCUSSION: CTS clustering methods are highly heterogeneous in assumptions, capabilities, and scalability. Distance-based approaches dominate, but model-based methods offer richer modeling potential, while feature-based ones emphasize performance at the cost of flexibility. CONCLUSION: This review highlights methodological diversity and gaps in CTS clustering. The proposed typology and Web application aim to help researchers choose appropriate methods to choose appropriate methods for their data.