Semantic feature extraction from technical texts with limited human intervention
Rajeev Agarwal · 1995
Natural Language Processing (NLP) and message systems often use semantic information in order to perform lexical and syntactic disambiguation and to assist them in understanding the text. Such information is domain-specific in nature and hence difficult to acquire in an automatic manner. This causes a problem whenever an NLP system is moved from one domain to another. Portability of an NLP system can be improved if these semantic features can be acquired with limited human intervention. The semantic information needed by an NLP system may take several different forms. This dissertation focuses on two such semantic features--semantic classes present in a given domain, and lexico-semantic patterns that exist between content words in the domain. This document discusses the techniques that are used to extract these semantic features from a domain with limited human intervention. Semantic classes are discovered by clustering different objects on the basis of the lexico-syntactic environments in which they appear in the corpus. The results of some experiments with augmenting the noun semantic classes with class information obtained from WordNet are presented. A methodology for formally evaluating the semantic classes extracted by the system against classes provided by experts is also presented. Once semantic classes have been obtained, they are then used to generate lexico-semantic patterns that are prevalent in the given domain. A noteworthy feature of this research is that the techniques used to acquire the semantic features require very limited human intervention. The combination of distributional and taxonomic techniques to obtain a set of semantic classes for a given domain has also been found to be useful.