Identifying contextual information for multi-word term extraction
Diana Maynard, S Ananiadou · Research Explorer (The University of Manchester) · 1999
Methods for multi-word term extraction have traditionally involved statistical techniques. More recently, hybrid techniques have been evolving which incorporate some linguistic knowledge. This information is generally very shallow, and researchers have tended to ignore any real understanding of either terms or the context in which they appear. We adopt an approach which uses a variety of knowledge sources - syntactic, semantic and statistical - and attempt to both enlighten and make use of the theoretical foundations of terminology in a practical application. We seek particularly to identify those parts of the context which are most relevant to the terms. We incorporate contextual weights, based on a new similarity measure, into a method for term recognition, and thereby improve the ranking of terms and enable disambiguation to be achieved. 1 Introduction Techniques for term extraction are becoming increasingly important, due to the existence of large volumes of textual data, coupled ...