That's Nice … What Can You Do With It?
Anja Belz · Computational Linguistics · 2009
A regular fixture on the mid 1990s international research seminar circuit was the "billion-neuron artificial brain" talk.The idea behind this project was simple: in order to create artificial intelligence, what was needed first of all was a very large artificial brain; if a big enough set of interconnected modules of neurons could be implemented, then it would be possible to evolve mammalian-level behavior with current computationalneuron technology.The talk included progress reports on the current size of the artificial brain, its structure, "update rate," and power consumption, and explained how intelligent behavior was going to develop by mechanisms simulating biological evolution.What the talk didn't mention was what kind of functionality the team had so far managed to evolve, and so the first comment at the end of the talk was inevitably "nice work, but have you actually done anything with the brain yet?" 1 In human language technology (HLT) research, we currently report a range of evaluation scores that measure and assess various aspects of systems, in particular the similarity of their outputs to samples of human language or to human-produced goldstandard annotations, but are we leaving ourselves open to the same question as the billion-neuron artificial brain researchers? Shrinking HorizonsHLT evaluation has a long history.Spärck Jones's Information Retrieval Experiment (1981) already had two decades of IR evaluation history to look back on.It provides a fairly comprehensive snapshot of HLT evaluation at the time, as much of HLT evaluation research was in the field of IR.One thing that is striking from today's perspective is the rich diversity of evaluation paradigms-user-oriented and developer-oriented, intrinsic and extrinsic 2 -that were being investigated and discussed on an equal footing