Almost Total Recall: Semantic Category Disambiguation Using Large Lexical Resources and Approximate String Matching

Pontus Stenetorp, Sampo Pyysalo, Sophia Ananiadou, Jun'ich Tsujii · Research Explorer (The University of Manchester) · 2011

In this study we investigate how semantic category disambiguation can be used to sup-port other Natural Language Processing tasks and annotation efforts. While previous re-search has mostly cast semantic category dis-ambiguation purely as a classification task, we propose a task setting analogous to dynamic width beam search that allows for multiple se-mantic categories to be suggested while aim-ing to minimise the number of suggestions and maintain high recall. We base our ap-proach on a recent machine learning-based system and evaluate it on six recently intro-duced corpora, one incorporating as many as 17 semantic categories, our system performs in the recall range of 98.6 % to 99.5 % while keeping the average number of semantic cate-gories proposed in the range of 1.3 to 2.0. The level of performance suggests that the sys-tem is adequate to meet the human require-ments of human annotators and could suc-cessfully be used for annotation support. The introduced system and all related resources are freely available for research purposes at:

Read the paper · More papers on PaperTik