Estimating the True Performance of Classification-Based NLP Technology

James R. Nolan · 1997

Many of the tasks associated with natural language processing (NLP) can be viewed as classification problems. Examples are the computer grading of student writing samples and speech recognition systems. If we accept this view, then the objective of learning classifications from sample text is to classify and predict successfully on new text. While success in the marketplace can be said to be the ultimate test of validation for NLP systems, this success is not likely to be achieved unless appropriate techniques are used to validate the prototype. This paper discusses useful validation techniques for classification-based NLP systems and how these techniques may be used to estimate the true performance of the system.

Read the paper · More papers on PaperTik