Evaluation of Machine Translations by Reading Comprehension Tests and Subjective Judgments

Sheila M. Pfafflin · Mech. Transl. Comput. Linguistics · 1965

This paper discusses the results of an experiment designed to test the quality of translations, in which human subjects were presented with IBM-produced machine translations of several passages taken from the Russian electrical engineering journal Elektrosviaz, and with human translations of some other passages taken from Telecommunications, the English translation of Elektrosviaz. The subjects were tested for comprehension of the passages read, and were also asked to judge the clarity of individual sentences. Although the human translations generally gave better results than the machine translations, the differences were frequently not significant. Most subjects regarded the machine translations as comprehensible and clear enough to indicate whether a more polished human translation was desirable. The reading comprehension test and the judgment of clarity test were found to give more consistent results than an earlier procedure for evaluating translations, since the questions asked in the current series of tests were more precise and limited in scope than those in the earlier scries. In view of the considerable effort currently going into mechanical translation, it would be desirable to have some way of evaluating the results of various translation methods. An individual who wishes to form his own opinion of such translations can, of course, read a sample, but this procedure is unsatisfactory for many purposes. To indicate only one difficulty, individuals vary widely in their reactions to the same sample of translation. However, a previous attempt by Miller and Beebecenter 1 to develop a more satisfactory approach gave discouraging results. When ratings of the quality of passages were used, it was found that subjects had considerable difficulty in performing the task, and were highly variable in their ratings; while information measures, which were also used, proved very time-consuming. Furthermore, neither of these methods provided a direct test of the subject's understanding of the translated material. The present study explored two other approaches to the evaluation problem, namely, reading comprehension tests, and judgments of the clarity of meaning of individual sentences. The approach through testing of reading comprehension provides a direct test of at least one aspect of the quality of translation. Judgments of sentence clarity do not, but they are likely to be simpler to prepare and may have applicability to a wider range of material. Both types of tests might therefore be useful for different evaluation problems if they proved to be effective. While the previous results with a rat* The author wishes to express her appreciation to Mrs. A. Werner, who prepared translations for preliminary tests and advised on preparation of the final text, and to D. B. Robinson, Jr., L. Rosier, and B. J. Kinsberg for their contributions to selection of passages and preparation of questions used in the reading comprehension tests. ing technique are not encouraging for a judgment method, the assignment of one rating of over-all quality to a passage is a fairly complex task. We hoped that by asking subjects to judge sentences rather than passages, and to judge for clarity of meaning only, rather than quality generally, the subjects' task would be simplified and the results made more reliable.

Read the paper · More papers on PaperTik