Web data cleansing and preparation for ontology extraction using WordNet

Keng-Woei Tan, Hyoil Han, Ramez A. Elmasri · 2002

The explosive growth of data on the World Wide Web makes information management and knowledge discovery increasingly difficult. Applying database techniques to manage Web information can help in solving these problems. One difficulty encountered is that Web documents, unlike structured databases, contain unstructured and semi-structured data. Our hypothesis is that creating ontologies to describe the semantics of Web data is the key to bridging the gap between semi-structured data and structured databases, and hence to facilitating the application of database techniques. We extract an ontology (or conceptual schema) from a set of Web pages in a particular application domain automatically. The prototype we are constructing is called WebOntEx (Web Ontology Extraction). This paper describes the data preparation process and the semantic resolution process of the WebOntEx project to build a meta-database and a Web database.

Read the paper · More papers on PaperTik