Error-Annotated Corpus of Latvian

Deksne Daiga, Skadi ncedil a Inguna · Frontiers in artificial intelligence and applications · 2014

This paper reports on the development of the annotated Latvian language error corpus designed for grammar checker development and evaluation. We describe the error classification system introduced for this purpose, the annotation process, and guidelines. Two corpora (the corpus of student papers and the balanced text corpus) consisting of a total of 20,877 sentences have been created and annotated. A general characterisation of the corpora and a summary of the annotation results are presented.

Read the paper · More papers on PaperTik