An Improved Random Forest Algorithm Based on Spark
Machine Learning Theory and Practice · 2022
Classification algorithms are an important branch of data mining, and are also of great importance in the era of BD.The random forest algorithm (RFA) is one of the classification algorithms and is widely used in various industries for its good classification performance.However, the performance of RFA is not so good when dealing with high-dimensional data and unbalanced data.The main objective of this paper is to improve the RFA based on Spark.In this paper, we read a large amount of relevant algorithm literature in terms of algorithm research, and gain a comprehensive understanding of what feature selection and unbalanced classification are, as well as what characteristics they have and how these problems should be solved.It then focuses on how some domestic and international scholars have solved these problems.This paper focuses on studying and analysing the strengths and weaknesses of the RFA, and makes relevant improvements to address the two weaknesses of the RFA.In order to solve the problems of the RFA in the field of feature selection and the field of unbalanced classification, the optimization is improved respectively, and the parallelized design of the optimized algorithm is finally implemented on Spark.