Hybrid data mining algorithm in cloud computing using MapReduce framework
Siddharth Sahay, Suruchi Khetarpal, Tribikram Pradhan · 2016
'Data mining' has transformed into a ubiquitous term in the world of IT and Computer Science in recent times. Developments in this field have been countless. Using one of Apriori algorithm's numerous variants with a couple of insightful additions can significantly improve upon the existing standard of Data Mining. In this paper a new approach to considerably reduce the time complexity of the database scan has been proposed. This has been achieved by using the MapReduce framework for Hadoop Distributed File System (HDFS). Coupled with Cloud computing, which handles large data sets and processing remotely, the resultant system - that uses MapReduce for the full table scan, the Pincer-Search Algorithm, and Cloud Computing - is a force to reckon with.