An Analysis of Errors in Google Neural Machine Translation

Yimeng Wang, Siqi Liu, Yingwei Ma · 2021 2nd International Conference on Information Science and Education (ICISE-IE) · 2021

Google Neural Machine Translation employs an encoder-decoder framework with an attention mechanism and it has resulted in a tremendous increase in efficiency. However, the quality of its translations is not yet comparable to that of human translations. In order to discover the specific problem of GNMT, the study analyzes the accuracy rate and types of errors in Google Neural Machine Translation by means of case studies. Based on the data collected, it is found that errors at the lexical level are most significant. The sparse data for certain terms in the Google Translate database, such as buzzwords, make undertranslation, a subcategory at the lexical level, the most significant translation error. These results proved that the users and developers may pay greater attention to the lexical level and try to match the selflearning capability of machine translation to the speed of partial vocabulary updates.

Read the paper · More papers on PaperTik