Performance of Web Information Gathering System using Ontology and Information Extraction

Y. Ramesh Kumar, U. Kartheek, Chandra Patnaik, Gannavarapu Ananth Kumar · 2012

Traditional methods on Information Extraction (IE) have focused on the use of supervised learning techniques such as hidden Markov models, self-supervised methods, rule learning, and Conditional Random Fields (CRF). WebNLP framework is based on CRF’s and markov models. These techniques learn a language model or a set of rules from a set of hand-tagged training documents and then apply the model or rules to new texts. Models learned in this manner are effective on documents similar to the set of training documents, but extract quite poorly when applied to documents with a different genre or style which is usually found on web. As a result, this approach has difficulty scaling to the Web due to the diversity of text styles and genres on the Web and the prohibitive cost of creating an equally diverse set of hand tagged documents. In this paper we propose an adaptive IE system which uses Ontology Based Information Extraction techniques that extracts all relations by learning a set of lexico-syntactic patterns unlike WebNLP. It permits greater machine interpretability of content than that supported by XML, RDF and RDF Schema (RDF-S), by providing additional vocabulary along with a formal semantics. So, ontologies represent an ideal knowledge background in which to base text understanding and enable the extraction of relevant information.

Read the paper · More papers on PaperTik