Extracting Knowledge Bases from table-structured Web Resources applied to the semantic based Requirements Engineering Methodology SoftWiki.
Rubén Navarro Piris · 2010
A lot of information on the Web is provided as HTML formatted tables and CSV1 files. Such tables contain semantic information that can be derived from the embedded environment of the table as well from the heading of each column. Often the problem of integrating and linking this information into semantic web applications occurs. One way to solve this is a transformation of these tables into OWL ontologies. The requirements engineering tool SoftWiki [DAR06] is an example of such a semantic web application. The SoftWiki methodology [LR09], based on a set of tools including the SoftWiki tool, is used to interpret and manipulate data based on the SoftWiki vocabulary. A common task using the SoftWiki methodology is to import data. The developed OntoWiki extension imports the issue tracker of Google Code2 projects which is provided as a CSV file and transforms it into an ontology. The paper describes a methodology that is divided into three parts: (1) Analysis of the table structure, (2) Analysis of the vocabulary and (3) Definition of a N33 template. In a typical table the structure is always the same: the columns contain the attributes and the rows the entities. The first row is the one that determines the semantic of the other rows, which contain data that follow this model. Because of its simplicity it is very space efficient, but it’s difficult to further analyze or use this information because of the lack of explicit semantic relations. city country Inhabitants universities