Concepts and tools for the effective and efficient use of web archives

Helge Holzmann · Institutional Repository of Leibniz Universität Hannover (Leibniz Universität Hannover) · 2019

Web archives constitute valuable sources for researchers in various disciplines. However, their sheer size, the typically broad scope and their temporal dimension make them difficult to work with. We have identified three views to access and explore Web archives from different perspectives: user-, data- and graph-centric. The natural way to look at the information in a Web archive is through a Web browser, just like the live Web is consumed. This is what we consider the user-centric view. The most commonly used tool to access a Web archive this way is the Wayback Machine, the Internet Archive's replay tool to render archived webpages. To facilitate the discovery of a page if the URL or timestamp of interest is unknown, we propose effective approaches to search Web archives by keyword with a temporal dimension through social bookmarks and labeled hyperlinks. Another way for users to find and access archived pages is past information on the current Web that is linked to the corresponding evidence in a Web archive. A presented tool for this purpose ensures coherent archived states of webpages related to a common object as rich temporal representations to be referenced and shared. Besides accessing a Web archive by closely reading individual pages like users do, distant reading methods enable analyzing archival collections at scale. This data-centric view enables analysis of the Web and its dynamics itself as well as the contents of archived pages. We address both angles: 1. by presenting a retrospective analysis of crawl metadata on the size, age and growth of a Web dataset, 2. by proposing a programming framework for efficiently processing archival collections. ArchiveSpark operates on standard formats to build research corpora from Web archives and facilitates the process of filtering as well as data extraction and derivation at scale. The third perspective is what we call the graph-centric view. Here, websites, pages or extracted facts are considered nodes in a graph. Links among pages or the extracted information are represented by edges in the graph. This structural perspective conveys an overview of the holdings and connections between contained resources and information. While this enables novel concepts of exploring Web archives, it also raises new challenges. We present the latest achievements in all three views as well as synergies among them. For instance, important websites that can be identified from the graph-centric perspective may be of particular interest for the users of a Web archive. The data-centric view is used in both ways, it benefits from the graph-centric view to guide data studies but is also employed to prepare the data for the other views, like extracting graphs from archival collections. Finally, by considering the three views as different zoom levels of the same Web archive, they can be integrated in a holistic data analysis pipeline.

Read the paper · More papers on PaperTik