Çok boyutlu veriden rastgele ormanlar ile düşük karmaşıklıklı güdümsüz öğrenme
Ali Selman Aydin, Tank Arici, Ahmet Bulut · 2014
With the ever increasing rate of digital information available from online sources, information has gone from being scarce to being abundant. Big data analytics require low complexity and distributed computing techniques. We propose the use of randomized decision trees and their ensemble in the form of a forest for unsupervised learning. Random probing of good attributes reduces the computational complexity making learning feasible on high-dimensional big data. Using an ensemble of trees improves the learning. We propose a new splitting measure for tree construction and an aggregation mechanism for predictive learning (unsupervised classification). The experiments on standard datasets show that our proposed proposed method outperforms the state-of-the-art.