Evaluating parts-of-speech taggers for use in a text-to-scene conversion system
Kevin Glass, Shaun Bangay · 2005
This paper presents parts-of-speech tagging as a rst step towards an autonomous text-to-scene conversion system. It categorizes some freely available taggers, according to the techniques used by each in order to automatically identify word-classes. In addition, the performance of each identi ed tagger is veri ed experimentally. The SUSANNE corpus is used for testing and reveals the complexity of working with di erent tagsets, resulting in substantially lower accuracies in our tests than in those reported by the developers of each tagger. The taggers are then grouped to form a voting system to attempt to raise accuracies, but in no cases do the combined results improve upon the individual accuracies. Additionally a new metric, agreement, is tentatively proposed as an indication of con dence in the output of a group of taggers where such output cannot be validated.