Challenges in Information Extraction from Text for Knowledge Management

Fabio Ciravegna · 2001

Nowadays large part of knowledge is stored in an unstructured textual format. Texts cannot be queried in simple ways and therefore the contained knowledge can neither be used by automatic systems, nor be easily managed by humans. The traditional process of manual knowledge identification and extraction by knowledge engineers used in KM is a complex time consuming process, as it requires a great deal of manual input. An example of such process is the collection of interviews to experts (‘protocols’) and their analysis by knowledge engineers in order to codify, model and extract the knowledge of an expert in a particular domain. In this context, Information Extraction from texts (IE) is one of the most promising areas of HLT for Knowledge Management (KM) applications. IE is an automatic method for locating important facts in electronic documents, e.g. information highlighting for document enrichment or for information storing for further use (such as populating an ontology with instances). IE as defined above is the perfect support for Knowledge Identification and Extraction as it can – for example provide support in protocol analysis either in an automatic way (unsupervised extraction of information) or semi-automatic way (e.g. helping knowledge engineers locating the important facts in protocols, via information highlighting). It is widely agreed that the main barrier to the use of IE is the difficulty in adapting IE systems to new scenarios and tasks, as most of the current technology still requires the intervention of IE experts. This makes IE a technology difficult to apply, because personnel skilled in IE are difficult to find in industry, especially in small medium enterprises [Ciravegna 2000]. A main challenge for IE for the next years is to enable people with knowledge of Artificial Intelligence (e.g. knowledge engineers) but no or scarce preparation on IE and Computational Linguistics to build new applications/cover new domains. This is particularly important for KM: IE is just one of the many technologies to be used in building complex applications: wider acceptance of IE will come only when IE tools will not require any specific skill apart from notions of KM. A number of Machine Learning based tools and methodologies are emerging [Freitag 1999, Ciravegna 2001], but the road to fully adaptable and effective IE systems is still long. In this paper, I will focus on two main challenges for adaptivity in IE for KM that in my opinion are paramount in the current scenario: 1. Automatic adaptation to different text types 2. Human-centred issues in copying with real users.

Read the paper · More papers on PaperTik