Modelling the syntactic contextual information for term extraction
Roberto Basili, Mt Pazienza, Fm Zanzotto · Cineca Institutional Research Information System (Tor Vergata University) · 2001
Terms are key components of terminological knowledge bases (TKBs).These are valuable but very expensive resources for a wide range of applications devoted to knowledge access and management.A large number of approaches to corpus-driven term knowledge base acquisition and extension have been proposed.Syntactic constraints are widely used to characterize terms in source texts.However, contextual information is generally exploited for semi-automatic acquisition of semantic relations among terms and a rich notion of semantic context is usually adopted.The aim of our research is to investigate whether syntactic context (i.e.structural information on local term contexts) can be used for determining "termhood" of given term candidates.A weakly supervised model is here proposed where predictive rules are built over the grammatical representation of the contexts available from limited terminological resources.Extensive experimental evidence derived from the analysis of a large legal corpus and a controlled terminology suggests the viability of this automatic method over unrestricted texts.