An efficient similarity measurement method and its classification effect

Hui YUAN, Zhanglu Tan, Fuhao Wang · Scientia Sinica Technologica · 2022

High-dimensional data classification is extremely significant in statistical analysis. However, the classification method still suffers from high noise sensitivity, a large calculation amount, and low accuracy because of the measurement distance it relies on, resulting in poor classification results. We propose an improved similarity measurement method, which we used to improve the classification effect, aiming to address the low accuracy and efficiency of high-dimensional time-series data classification. We propose a similarity search algorithm based on Euclidean distance and a 1-nearest neighbor (1-NN) classification technical framework. First, we used discrete wavelet transform (DWT) to decompose and reconstruct the sequence. Thereafter, we developed a local high-frequency DWT method to achieve dimensionality reduction and noise reduction. We combined the concepts of volatility and rank correlation coefficient based on the distance function and improved relative deviation and volatility trend consistency. The experimental results based on 40 UCR time-series datasets revealed that the 1-NN classification accuracy method proposed in this paper is superior to the dynamic time warping, FastDTW, and longest common subsequence measurement methods, with a confidence level of more than 85%. It also confirmed that the accuracy and speed of the 1-NN classification framework significantly improved. The findings of this research add to the theoretical basis of similarity measurement and serve as a valuable reference for data mining applications in intelligent system management and time-series statistics.

Read the paper · More papers on PaperTik