Relation Extraction for Ontology Construction
Kate G. Byrne · 2006
This proposal is for a programme of work leading to the building of a data querying application. The starting point is a collection of relational databases holding cultural heritage material from the National Collections of Scotland. The data is a mixture of fixed fields and free text, supported by background material such as domain thesauri. The goal is to produce a system for running queries against this material that does not assume the user has expert knowledge of the data structure or the specialist domain terminology. Two core tasks are proposed as necessary steps towards the goal: • Extraction of two-place relations from free text: The purpose of this step is to translate key facts from the textual material into a standardised format. A combination of rule-based and machine learning approaches is planned. • Automatic assembly of all relevant data into an ontology: An ontology is defined here as a graph of two-place relations where the edges represent predicates and the nodes the entities they apply to. All of the relevant information — from database fields, domain thesauri and the extracted textual relations — will be combined into such a graph. The application will run against the populated ontology. An interactive interface is proposed, in which a tailored summary based on the user’s initial query is generated and the user is then invited to refine the query based on this information. Evaluation of each subtask is planned, and the overall criteria for success will be that query performance is comparable to what is currently available in terms of speed and range, whilst also producing improved results for the non-expert user.