SparkFSM: A Highly Scalable Frequent Subgraph Mining Approach using Apache Spark
Bismita Jena, Cynthia Khan, Rajshekhar Sunderraman · 2018
Knowledge mining from graph data has attracted many researchers over the past several years. With the evolution of internet, computer technology, social networking sites, and web logs, graphs have become a very crucial dataset for mining and finding appropriate knowledge. Based on the application, graphs take different forms, such as airline flight information graph is mostly directed and smaller graphs, chemical compound structures are small and undirected, and social network graphs are very large single graphs based on the different types of associations between people. Earlier, our first attempt to use Hadoop (MapReduce model) to mine the directed frequent subgraphs from a large group of smaller graphs, which is famously known as transaction graphs, proved to be very scalable over the memory-based or database-oriented approaches. In this paper, we introduce SparkFSM, which not only handles undirected and directed graphs, but is also very scalable and efficient in handling isomorphism with the relatively new technology in industry (Spark/Scala). The combination of Spark with the functional style language Scala has established to be a de Facto standard while dealing with in-memory large data processing.