A Time Series Data Augmentation Method based on SMOTE
Hongchun Qu, Zheng Zhang · 2024
Neural networks have become a focal point in the realms of time series analysis and data mining. However, unlike in the field of images, time series datasets are typically smaller in scale and exhibit substantial imbalance in data types, making models prone to overfitting. The conventional SMOTE (Synthetic Minority Over-sampling Technique) method is widely used for addressing imbalanced datasets through data augmentation. Nevertheless, when applied to time series data, the classical SMOTE method struggles to effectively accommodate the characteristics of time sequences, resulting in suboptimal performance on time series datasets. To address this issue, this paper proposes an enhanced algorithm, namely M-SMOTE. This algorithm relies on specific partitioning rules to classify minority class samples into safe points and noise points. Subsequently, linear interpolation is applied only between safe points and the centroid of minority class samples (i.e., the mean of minority class samples), effectively mitigating issues related to data marginalization. Additionally, the M-SMOTE algorithm avoids the uncertainty associated with selecting the k value in traditional SMOTE algorithms. In situations where time series with similar features may undergo phase shifts, the method employs dynamic time warping to obtain more reasonable similarities between different time sequences. Comparative experiments are conducted using deep convolutional neural networks (CNNs) and recurrent neural networks (RNNs). The results demonstrate a significant improvement in the performance of the mentioned models when trained on datasets processed using the M-SMOTE method. This approach provides an effective and innovative solution for addressing imbalance issues in time series datasets.