Methods and systems for ontology learning, exploitation, and analysis
Jiang Xing · 2010
While keyword based techniques continue to be the most popular option for information services, the limitations inherent in keywords routinely generate unsatisfactory results.As a promising alternative, ontology based solutions have been proposed to provide effective information services by exploiting ontologies for representing and organizing information.This thesis addresses the key issues in adopting ontology based solutions by presenting a collection of methods and systems for ontology building, ontology exploitation, and ontology analysis.In any ontology based solution, ontologies firstly have to be created for representing and organizing information.However, ontology building is well known to be a tedious process.Manually acquiring knowledge for building domain ontologies requires much time and resources.To ease the efforts of building ontologies, we develop a system called Concept-Relation-Concept Tuple based Ontology Learning (CRCTOL) for automatically learning ontologies from domain specific text documents.By using a full text parsing technique and incorporating both statistical and lexico-syntactic methods, the ontologies learned by our system are more concise and contain a richer semantics in terms of the range and number of semantic relations compared with alternative systems.To provide ontology assisted services, we study a major application of ontology based solutions, namely ontology based information retrieval.In view of the limitations of the existing user models, we develop an ontology based user model, called user ontology, which is a specialization of the domain ontology by assigning each concept and relation of the domain ontology with a specific value for indicating a user's interests.We have developed methods for learning and inferencing in the user ontology model and integrated it into a semantic search engine called OntoSearch for providing personalized document retrieval.The experimental results, based on the ACM digital library and the Google i