Workshop on the evaluation of natural language processing systems
Martha Stone Palmer, Tim Finin · 1990
In the past few years, the computational linguistics research community has begun to wrestle with the problem of howtoevaluate its progress in developing natural language processing systems. With the exception of natural language interfaces there are few working systems in existence, and they tend to focus on very di#erent tasks using equally di#erent techniques. There has been little agreement in the #eld about training sets and test sets, or about clearly de#ned subsets of problems that constitute standards for di#erent levels of performance. Even those groups that have attempted a measure of self-evaluation have often been reduced to discussing a system's performance in isolation - comparing its current performance to its previous performance rather than to another system. As this technology begins to move slowly into the marketplace, the lack of useful evaluation techniques is becoming more and more painfully obvious.