System for Post-Editing and Automatic Error Classification of Machine Translation
Jozef Kapusta, Daša Munková, Martin Drlík · 2016
We describe a system for manual and automatic classification of errors in machine translation output. The system was designed based on Vilar´s at al. (2006) error classification and TAUS error typology (2013) taking into consideration the target language- Slovak language. Slovak is an inflectional language with a “rich morphology” and belongs to the synthetic languages with a loose word order. System features manual evaluation in the form of post-editing and the detection and classification of the basic errors (Language, Accuracy, Terminology or Style) done by post-editors. Beside the manual classification of errors, based on POS tagging or morphological tag annotation, it features also the automatic error detection and classification into morphological errors (Incorrect word form), reordering errors (Word order), missing words, lexical errors (Incorrect words),unknown words and punctuation. The aim of this paper is to proposed framework and design a system for the purpose to determine which error types cause a problem for a given machine translation system when translating a text of various styles (administrative, scientific and publicistic) into inflectional language as Slovak is. Normal 0 21 false false false SK JA X-NONE /* Style Definitions */ table.MsoNormalTable {mso-style-name:Normalna tabuľka; mso-tstyle-rowband-size:0; mso-tstyle-colband-size:0; mso-style-noshow:yes; mso-style-priority:99; mso-style-parent:; mso-padding-alt:0cm 5.4pt 0cm 5.4pt; mso-para-margin:0cm; mso-para-margin-bottom:.0001pt; mso-pagination:widow-orphan; font-size:10.0pt; font-family:Calibri,sans-serif;}