Research on Large-Scale Data Mining and Machine Learning Model for Accurate Prediction
Min Chen, Shang Yu, Yongsheng Deng · 2025
Traditional forecasting methods are often unable to meet the needs of accurate forecasting when faced with large-scale and high-dimensional data. Considering the diversity and complexity of data, as well as the continuity or classification of prediction targets, this paper chooses the random forest (RF) algorithm as the basic model. By assembling multiple decision trees and aggregating their outputs, the RF algorithm enhances prediction accuracy and demonstrates robust resistance to overfitting. In regression tasks, RF calculates the mean of predictions from individual trees to produce its final output; for classification tasks, it employs a majority voting scheme to determine the outcome class. Regarding missing data and standardization, this study adopts the mean or median for imputing missing numerical features and the mode for categorical features. Additionally, Z-score normalization is applied to transform the data into a distribution characterized by a mean of 0 and a standard deviation of 1. The empirical results indicate that the approach taken in this research exhibits strong generalizability and predictive precision when managing extensive datasets and intricate feature interactions. Compared with the traditional linear regression and support vector machine (SVM) model, RF model shows a higher R value in regression problems, showing a strong explanatory power. In the classification problem, although the precision is high, the recall rate and F1 value are relatively low, indicating that it is lacking in identifying positive cases. The research in this paper not only helps to improve the accuracy of prediction and optimize the decision-making process, but also provides new ideas and methods for the application of data mining and machine learning in accurate prediction.