Term sense disambiguation using a domain-specific thesaurus
Diana Maynard, Sophia Ananiadou · 1998
Term extraction is important for many information systems applications. Although terms should be monoreferential, in reality they exhibit a high degree of ambiguity. This paper describes a method for automatic term sense disambiguation based on the identification of relevant contextual information from sublanguage corpora. We combine basic semantic roles derived from the corpus with domain-specific semantic categories from UMLS, a specialised medical thesaurus, and use an EBMT-based matching algorithm to compare related terms. 1 Introduction The increasing volumes of electronic textual data and the development of new applications have given rise to a variety of disambiguation techniques. Term ambiguity can greatly reduce the efficiency of term extraction applications such as information extraction and retrieval, dictionary construction, automatic indexing, machine translation and hypertext linking. Traditional methods of sense disambiguation involve comparing the distribution of the w...