Term Extraction Through Unithood and Termhood Unification
Thuy Vu, Ai Ti Aw, Min Zhang · 2008
Term Extraction (TE) is an important component of many NLP applications. In general, terms are extracted for a given text collection based on global context and frequency analysis on words/phrases association. These extracted terms represent effectively the text content of the collection for knowledge elicitation tasks. However, they fail to dictate the local contextual information for each document effectively. In this paper, we refine the state-of-the-art C/NC-Value term weighting method by considering both termhood and unithood measures, and use the former extracted terms to direct the local term extraction for each document. We performed the experiments on Straits Times year 2006 corpus and evaluated our performance using Wikipedia termbank. The experiments showed that our model outperforms C/NC-Value method for global term extraction by 24.4 % based on term ranking. The precision for local term extraction improves by 12 % when compared to pure linguistic based extraction method. 1