The Light-Weight Semantic Web: Integrating Information Extraction and Information Retrieval for Heterogeneous Environments
Jens Graupmann, Ralf Schenkel · 2005
Today’s Web, large intranets and even the documents collected by a single user are enormous sources of distributed, heterogeneous information that cannot be easily mastered. Syntactical and semantical differences as well as missing semantic annotations make effective query evaluation on such corpora a hard task. The Semantic Web aims at providing a standard for semantic annotations, but has not yet made large progress in the real world. This paper presents a light-weight version of the Semantic Web. We advocate the use of Information Extraction tools to automatically detect and annotate important classes of information that are frequently used in queries, like locations and dates. We propose a query language that can exploit the extra annotations and allows novel range and join conditions.