Towards a Semantic Information Extraction Approach from Unstructured Documents.
Massimo Ruffolo, Lorenzo Gallucci, Nicola Leone, Marco Manna, Domenico Saccà · SEBD · 2006
Recognizing and extracting meaningful information from semiand unstructured documents, taking into account their semantics, and storing them into database is an important problem in the context of information access and retrieval. This paper describes a novel logic-based approach to information extraction from both semiand unstructured documents. The approach, implemented in the HiLeX system, is founded on a new two-dimensional representation of documents constituting a unified abstract representation of both HTML pages and flat text documents. The semantics of information to be extracted is encoded by means of ontology expressed in DLP, an extension of disjunctive logic programming for ontology representation and reasoning, which has been recently implemented on top of the DLV system. Unlike previous systems, which are mainly syntactic, HiLeX combines both semantic and syntactic knowledge for a powerful information extraction. Each extraction pattern belongs to an ontology class and is expressed using regular expressions and/or an ad hoc two-dimensional language exploiting the document two-dimensional representation. The execution of DLP reasoning modules, encoding the HiLeX language expressions in term of logic rules, yields the actual extraction of information from the input document. HiLeX allows the semantic information extraction from both HTML pages and flat text documents by using synthetic and very expressive extraction patterns.