Detection of Weak Relevant Variables using Random Forests

Shuhei Kimura, Masato Tokuhisa · 2020

In general, the purpose of a machine learning technique is to obtain an approximation of a function based on a set of training data. To obtain a good approximation of the function, the input variables irrelevant to the output should be removed in advance. The process used to remove the irrelevant input variables is known as feature selection. The existing feature selection methods generally select input variables that make it possible to approximate the function more accurately. As a consequence, these methods often fail to detect input variables that weakly affect the output, especially when the training data given are insufficient. In some practical applications of machine learning techniques, however, all of the input variables that actually affect the output must be detected. In this study, we propose a new feature selection method to overcome this drawback of the existing methods. Our method evaluates the relevance of a certain input variable by comparing it with a random variable. We show, through numerical experiments, that the proposed method is capable of detecting even input variables that weakly affect the output.

Read the paper · More papers on PaperTik