Research on Time Series Data Prediction Based on Machine Learning Algorithms
Zhiwei Zheng, Yang Yang, Junyu Zhou, Fengxu Gu · 2024
This study focuses on small-sample data with four different distributions, introducing six feature variables. KNN, decision tree, random forest, multilayer perceptron (MLP), support vector regression (SVR), and ridge regression models are constructed to compare the predictive performance of these machine learning models on small-sample, multi-feature time series data. Firstly, the years are transformed into feature variables, converting temporal data into supervised data. After normalizing the feature variables, the datasets are split into training and testing sets in an 8:2 ratio. Subsequently, hyperparameter optimization is conducted for each model, involving the precise definition of optimal parameter ranges and determining the best hyperparameter configurations on the four datasets using a grid search algorithm. Finally, through comparing the predicted data on the testing set with the real data and utilizing visualization techniques, the study presents relevant results. The findings indicate that the ridge regression model demonstrates superior generalization capabilities, SVR is suitable for nonlinear data, ensemble learning (represented by random forests) outperforms single learners in generalization, and simple deep learning models (represented by MLP) have advantages in handling linear data. Future work aims to explore the performance of different machine learning models on large-sample, multi-feature time series data.