Statistical Tools for Corpus Analysis: A Tagger and Lemmatizer for Italian

Eugenio Picchi · 1994

We present the most recent addition to the PiSystem, an integrated set of tools for monoand bilingual corpus creation and manipulation and dictionary construction. The new component is a statistical part-of-speech tagger and lemmatizer. The methodology adopted resembles that of similar procedures for other languages but the PiTagger has been developed to meet the particular requirements of a highly inflected language such as Italian. Texts analysed by the PiTagger can then be directly interrogated using the tagged corpus query procedures included in the system. The philosophy behind a procedure for sense disambiguation now being designed and tested is also briefly described.

Read the paper · More papers on PaperTik