Implementation of Breiman's Random Forest Machine Learning Algorithm

Frederick Livingston · 2005

This research provides tools for exploring Breiman’s Random Forest algorithm. This paper will focus on the development, the verification, and the significance of variable importance. Introduction A classical machine learner is developed by collecting samples of data to represent the entire population. This data set is usually subdivided into two or more dataset. Part of the dataset set is commonly use for developing the machine learner, and the remaining data is use for evaluation. Often this data set is imbalanced; the data consists of only a very small minority of the data. Imbalanced machine learners tend to perform poorly with the classification of fraud detection, network intrusion, rare disease diagnosing, etc [1, 2]. This is due to imbalanced sampling during developing the machine learner. During the testing phase these rare cases are unseen during the training phase and are usually misclassified. Leo Breiman, a statistician from University of California at Berkeley, developed a machine learning algorithm to improve classification of diverse data using random sampling and attributes selection. This project involved the implementation of Breiman’s random forest algorithm into Weka. Weka is a data mining software in development by The University of Waikato. Many features of the random forest algorithm have yet to be implemented into this software.

Read the paper · More papers on PaperTik