Subgroup discovery on big data: Pruning the search space on exhaustive search algorithms
F. Padillo, José María Luna, Sebastián Ventura · 2016
Subgroup Discovery is a broadly applicable supervised local pattern mining method to search relations between different properties with respect to a target variable. With the exponential growth in data storage, the massive data gathered has hampered the performance of current techniques. In this regard, our aim is to propose two new algorithms to discover subgroups on Big Data by using MapReduce. Apache Spark was used to tackle the Big Data requirements. The experimental study includes more than 50 large datasets and a set of efficient algorithms. Search spaces bigger than 1.276 · 1015subgroups are used. The experimental study reveals the alluring results in efficiency when optimistic estimates are considered, as well as demonstrating the usefulness of using Apache Spark to tackle Big Data.