Comparative Analysis of Web Scraping Methodologies Using Generative AI

M. Pushpalatha, Madhu Shree Aravindan · 2025

Web scraping is a method of extracting information from websites, and it plays a crucial role in data collection for various applications such as market research, academic studies, and competitive analysis. Traditional web scraping can be challenging, especially when dealing with multiple websites that may have different structures and content formats. The complexity of writing and maintaining scrap scripts for each site adds to the difficulty. To address these challenges, Generative AI (GenAI) offers a powerful solution for web scraping. This paper explores the use of Retrieval Augmented Generation (RAG) model architecture and ScrapeGraphAI, a Python library that leverages Large Language Models (LLMs), RAG and graph logic to automate the creation of scraping pipelines. ScrapeGraphAI employs LLMs to understand the desired information and build the necessary scraping logic. It supports multi-format data extraction, handles both single-page and multi-page scenarios, and generates scraping pipelines based on user instructions, significantly reducing manual coding efforts. The Gemini LLM is introduced as an alternative tool for web scraping. This study compares the performance of RAG, ScrapeGraphAI and Gemini LLM in collecting data from various websites to gather information on upcoming academic events. The comparison focuses on the efficiency, accuracy, and ease of use of both methods. Through this analysis, we aim to highlight the advantages and limitations of using AI driven web scraping technologies in streamlining data collection processes.

Read the paper · More papers on PaperTik