Scalable Web Crawling: Harnessing the Power of Hadoop MapReduce in a Distributed Framework

Anjali Chennupati, Bhamidipati Prahas, Bharadwaj Aaditya Ghali, Manju Venugopalan · 2023

This paper demonstrates the implementation of Distributed Web Crawling using the Hadoop MapReduce framework on a distributed system where multiple Virtual Machines are connected through a master-slave relationship with the Hadoop MapReduce Framework as the underlay. Distributed web crawling optimizes data retrieval by involving multiple machines, leading to faster crawling, fault tolerance, and scalability. The proposed approach harnesses the strength of Hadoop MapReduce which is powerful in handling web data at scale for various applications, including search engine indexing, data mining, and large-scale analytics. Furthermore, this robust, scalable, and high-performance solution for gathering information from the vast World Wide Web enables deeper insights and broader applications within the realm of big data. A performance experiment demonstrates the model's efficiency compared to sequential crawling which reports a reduction in execution time ranging from 60 to 75%.

Read the paper · More papers on PaperTik