Sublanguage Dependent Evaluation: Toward Predicting NLP performances

Gabriel Illouz · 2000

In Natural Language Processing (NLP) Evaluation, such as MUC (Hirshman, 1998), TREC (Harman, 1998), GRACE (Adda et al., 1997), SENSEVAL (Kilgariff, 1998), metrics on the performances, such as precision, recall, or f-measure are used.Nevertheless, performance results are often average measurements computed over the complete test.They do not give any clues about the system's robustness.We conceive evaluations being not only a processs to show how good the systems are on a given dataset, but also as an aid for choosing which system or approach to use to build a NLP application for a specific subset of the language.In this case, knowing which system performs better on average does not help us to find which is the best for a given subset of a language.As a matter of fact, this aspect of the reuse paradigm is rarely investigated in the litterature about workbenches especially designed to adapt quickly to new language resources, such as GATE (Cunningham, 1997), In the present article, the existing approaches which take into account language heterogeneity and offer methods to identify sublanguages are presented.Then we propose a new metric to assess robustness and we study the existence of a correlation between the performance variations observed for POS tagging and the different sublanguages identified in the Penn Tree Bank Corpus.The work we present here is a first step in the development of predictive evaluation methods, intended to propose new tools to help in determining in advance the range of performance that can be expected from a system on a given dataset.

Read the paper · More papers on PaperTik