K-Random Forests: a K-means style algorithm for Random Forest clustering
Manuele Bicego · 2019
In this paper we present a novel clustering approach based on Random Forests, a popular classification and regression technique whose usability in the clustering scenario has been investigated to a lesser extent. In the clustering context, the most used class of approaches is based on the exploitation of a single Random Forest to derive a proximity measure between points, to be used with any distance-based clustering technique. On the contrary, our scheme exploits a set of Random Forests, each one devoted to model one cluster, in a spirit similar to that of the mixture models approach. These Random Forests, which provide flexible cluster descriptors, are iteratively updated using a K-means-like clustering algorithm. The proposed scheme, which we call K-Random Forests (K-RF), has been evaluated on five datasets: the obtained results suggest that it represents a valid alternative to classic Random Forest clustering algorithms as well as to other established clustering approaches.