Part-of-Speech tagging strategy for MIDIA: a diachronic corpus of the Italian language

Claudio Iacobini, Aurelio De Rosa, Giovanna Schirato · Pisa University Press eBooks · 2014

English. The realization of MIDIA (a balanced diachronic corpus of written Italian texts ranging from the XIII to the first half of the XX c.) has raised the issue of developing a strategy for PoS tagging able to properly analyze texts from different textual genres belonging to a broad span of the history of the Italian language. The paper briefly describes the MIDIA corpus; it focuses on the improvements to the contemporary Italian parameter file of the PoS tagging program Tree Tagger, made to adapt the software to the analysis of a textual basis characterized by strong morpho-syntactic and lexical variation; and, finally, it outlines the reasons and the advantages of the strategies adopted. Italiano. La realizzazione di MIDIA (un corpus diacronico bilanciato di testi scritti dell'italiano dal XIII alla prima meta del XX secolo) ha posto il problema di elaborare una strategia di PoS tagging capace di analizzare adeguatamente testi appartenenti a diversi generi testuali e che si estendono lungo un ampio arco temporale della storia della lingua italiana. Il paper, dopo una breve descrizione del corpus MIDIA, si focalizza sui cambiamenti apportati al file dei parametri dell'italiano contemporaneo per il programma di PoS tagging Tree-Tagger al fine di renderlo adeguato all’analisi di una base testuale caratterizzata da una forte variazione morfosintattica e lessicale, e evidenzia le motivazioni e i vantaggi delle strategie adottate.

Read the paper · More papers on PaperTik