Exploring Large Language Models for Analyzing Changes in Web Archive Content: A Retrieval-Augmented Generation Approach
Jhon G. Botello, Lesley Frew, José J. Padilla, Michele C. Weigle · 2024
Websites typically display only their most recent content. However, the dynamic nature of the web leads to frequent updates and deletions. Web archives preserve snapshots of earlier versions for those interested in tracking changes over time. Analyzing these changes often requires a manual process that relies on traditional methods focused on terms or phrase-level differences. This study explores the capability of Large Language Models (LLMs), specifically GPT-4o, through a Retrieval-Augmented Generation (RAG) approach for detecting changes in archived web pages. Using WARC-GPT, a RAG pipeline to interact with Web ARChive (WARC) files, we identify and analyze changes across a small set of U.S. federal environmental web pages that changed between 2016 and 2020. Our findings show that GPT-4o can effectively be used to detect inconsistencies in web archive content, including consideration of the change and the semantic context upon which the changes occurred. Our exploration represents an initial step toward using Artificial Intelligence (AI) for deeper and scalable web change analysis.