Distributed Multi-class Rule Based Classification Using RIPPER
Aruna Govada, Varsha S. Thomas, Ipsita Priyadarsini Samal, Sanjay K. Sahay · 2016
Traditional data mining (DM) has certain challenges viz. Scalability, high dimensionality, distributed data and often it also requires huge amount of computational resources in terms of space and time to extract the hidden patterns in the data. In addition, the data has to be available at one location. But in today's era the data are often inherently distributed in several databases. Hence, due to the limited bandwidth, centralized processing of the data is highly inefficient. Therefore, distributed computing becomes very important for efficient DM, both in terms of space and time. This can be done by developing a mechanism to mine the massive data by applying DM in a non-centralized way that distributes the work load seamlessly among the available sites. Therefore, in this paper, we propose an algorithm distributed multi-class rule based classification (DiRUC) which implements repeated incremental pruning to produce error reduction (RIPPER) at local level and then merges into a global level in a distributed manner. The algorithm first constructs the local rule sets for the distributed data and then at each iteration the local models are sent from one location to other. Finally, the global model is constructed by efficiently merging these local models and is made available at each site for further prediction of the class labels. The performance (accuracy and efficiency) analysis of the algorithm is done for the five data sets with different parameters and the result shows that the proposed approach DiRUC outperforms the normal RIPPER and Ischibuchi et al. island model.