Tag Disambiguation in Italian

Rodolfo Delmonte, Emanuele Pianta · 1999

In this paper we argue in favour of syntactically based tagging by presenting data from a study of a 1,000,000 word corpus of Italian. Most papers present approaches on tagging which are statistically based. None of the statistically based analyses, however, produce an accuracy level comparable to the one obtained by means of linguistic rules [1]. Of course their data are strictly referred to English, with the exception of [2, 3, 4]. As to Italian, we argue that purely statistically based approaches are inefficient basically due to great sparsity of tag distribution – 50 % only unambiguous tags. In addition, the level of homography is also very high: readings per word are 1.7 compared to 1.07 computed for English by [2] with a similar tagset. In a preliminary experiment we made we obtained 99,97 % accuracy in the training set and 99,03 % in the test set using syntactic disambiguation: data derived from statistical tagging is well below 95% even when referred to the training set. 1.

Read the paper · More papers on PaperTik