Cross-lingual tagger evaluation without test data
Żeljko Agić, Barbara Plank, Anders Søgaard · 2017
We address the challenge of cross-lingual POS tagger evaluation in absence of manually annotated test data.We put forth and evaluate two dictionary-based metrics.On the tasks of accuracy prediction and system ranking, we reveal that these metrics are reliable enough to approximate test set-based evaluation, and at the same time lean enough to support assessment for truly low-resource languages.