gSpan-H: An Iterative MapReduce Based Frequent Subgraph Mining Algorithm
M.H.Sangle · International journal of advance research and innovative ideas in education · 2016
In data mining applications, mining frequent subgraph from a large number of small graphs is an important operation. For extracting frequent subgraphs many algorithms have been proposed. But now a days, graph data grows both in size and quantity, therefore existing methods cannot extract frequent subgraph on a centralized machine. To overcome this some distributed solution using MapReduce is becoming important paradigm for computation on massive data. In experimented work, we investigate how to efficiently perform extraction of frequent subgraph over a large datasets using MapReduce. We propose a frequent subgraph algorithm called as gSpan-H which is iterative MapReduce based framework. This algorithm uses breadth first search strategy. This algorithm is isomorphism testing free approach for efficiently mine frequent subgraph. Our experiments with real life and large synthetic datasets validate the effectiveness of gSpan-H for mining frequent subgraphs from large distributed datasets.