Identification and Analysis of Post-Editing Patterns for MT
Declan Groves, Dag Schmidtke · 2009
For this work we have carried out a number of anal-ysis experiments comparing raw MT output pro-duced by Microsoft’s Treelet MT engine (Quirk et al., 2005) with its human post-edited counterpart, for English–German and English–French. Through these experiments we identify a number of interest-ing post-editing patterns, both textual (string-based) and constituent-based. In this paper we discuss our analysis methodologies, present some of our results and provide information on how this type of analy-sis can be of benefit to translation systems and post-editors, with a view to improving initial MT output and consequently post-editor productivity. In addi-tion, we also discuss the MT and post-editing work-flow at Microsoft and results from MT post-editing pilots for a number of different language pairs. 1