SFM-GMDH: Sparse Feature Mapping GMDH Network for Time Series Prediction
Shuaihu Liu, Heshan Wang · 2024
Time series prediction has become increasingly crucial in today’s big data context, providing a means to obtain future trends from historical observations. However, traditional time series prediction models are time-consuming and have many limitations when dealing with complex and large-scale datasets. These models often struggle to capture the potential deep correlations in the dataset, and their ability to accurately estimate future data is limited. In this context, the Group Method of Data Handling (GMDH) self-organizing network model has emerged. It effectively avoids the interference of personal subjective factors in the case of information explosion and incomplete information, and improves the accuracy of the model by continuously evolving to approximate polynomials. However, the independently applied GMDH network model has certain limitations when dealing with Chaotic time series datasets. To address this issue, this study introduces an innovative approach that combines sparse feature mapping module with GMDH in a network model. In this study, the input data is first processed by delay, and then represented the delayed sequence as random feature nodes through a random feature extractor. These random feature nodes were fed into the encoder network for encoding, in order to extract key features. Through the sparse representation of the generated vectors, the linear correlation between the newly generated feature nodes is effectively reduced. Finally, the sparse vector is fed into the GMDH network for more accurate prediction. To verify the superiority and robustness of the proposed network structure model, this study conducted extensive experiments using diverse time series datasets from the real world. The comparative experimental results show that the prediction model combining sparse feature mapping module and GMDH performs well in terms of performance and prediction error. This innovative method not only improves the predictive performance of the model, but also exhibits stronger adaptability and generalization when dealing with long time series data.