Web Data Commons – Extracting Structured Data from Two Large Web Corpora
Hannes Mühleisen, Christian Bizer · MADOC (University of Mannheim) · 2012
More and more websites embed structured data describing for instance products, people, organizations, places, events, resumes, and cooking recipes into their HTML pages using encoding standards such as Microformats, Microdatas and RDFa.The Web Data Commons project extracts all Microformat, Microdata and RDFa data from the Common Crawl web corpus, the largest and most up-todata web corpus that is currently available to the public, and provides the extracted data for download in the form of RDF-quads.In this paper, we give an overview of the project and present statistics about the popularity of the different encoding standards as well as the kinds of data that are published using each format.The embedded data is crawled together with the HTML pages by Google, Microsoft and Yahoo!, which use the the data to enrich their search results.These companies have so far been the only ones capable of providing insights into the amount as well as the types of data that are currently published on the Web using Microformats, RDFa and Microdata.While a previously published study by Yahoo!Research [4] provided many insight, the analyzed web corpus not publicly available.This prohibits further analysis and the figures provided in the study have to be taken at face value.However, the situation has changed with the advent of the Common Crawl.Common Crawl 1 is a non-profit foundation that collects data from web pages using crawler software and publishes this data.So far, the Common Crawl foundation has published two Web corpora, one dating 2009/2010 and one dating February 2012.Together the two corpora contain over 4.5 Billion web pages.