Normalization for Automated Metrics: English and Arabic Speech Translation
Sherri L. Condon, Gregory A. Sanders, Dan Parvaz, Alan Rubenstein, Christy Doran, John S. Aberdeen, Beatrice T. Oshika · 2009
tical Use (TRANSTAC) program has experimented with applying automated me-trics to speech translation dialogues. For trans-lations into English, BLEU, TER, and METEOR scores correlate well with human judgments, but scores for translation into Arabic correlate with human judgments less strongly. This paper provides evidence to sup-port the hypothesis that automated measures of Arabic are lower due to variation and in-flection in Arabic by demonstrating that nor-malization operations improve correlation between BLEU scores and Likert-type judg-ments of semantic adequacy — as well as be-tween BLEU scores and human judgments of the successful transfer of the meaning of indi-vidual content words from English to Arabic. 1