Dual-Metric Clustering for Multivariate Time Series: KMeans with DTW and QuadTree with Entropy
Samuel R. Torres, Raphael de Freitas Saldanha, Rocío Zorrilla, Vitor Ribeiro, Eduardo H. M. Pena, Fabio Andre Machado Porto · 2024
The efficacy of machine learning models are contingent on input data quality and model selection itself. In this work we highlight the importance of data quality, particularly in identifying regions within the input space that exhibit similar behavior. Clustering is used to group similar data, and is explored for their potential to enhance model performance by identifying these regions. The aim of this paper is to provide insights into the effectiveness of using clustering to improve machine learning model performance.