Knowledge extraction from big data using MapReduce-based Parallel-Reduct algorithm
Tapan Chowdhury, Susanta Kumar Chakraborty, Sanjit Kumar Setua · 2016
Extraction of knowledge and predictive analysis are the new challenges for the rapidly growing large volume of data to make the right decision at right time. It is difficult to store, analyze and visualize such large data volume with its diversities with standard data mining tools. Hence, in this paper, we develop a MapReduce approach of a Parallel-Reduct algorithm based on the rough set theory for knowledge extraction. It performs data and task parallelism with the help of Map and Reduce functions to find minimum reduct. An extensive experimental evaluation shows that the proposed MapReduce-based parallel approach effectively processes big data on the Hadoop platform and it is more efficient than the sequential approach to extract knowledge from large data sets under different coarseness.