Language Translation

A.F.R. Brown · Journal of the ACM · 1958

I must begin by admitting, as is scarcely necessary, that. I am at least several months away from being able to feed a new piece of French text into a computer and have an English translation come out at the other end. In trying out my basic idea, with verbally expressed rules on filing cards, I found it was possible to arrive in about 110 hours of work at a system that would translate 200 consecutive sentences from a French chemical journal into passable English. This was during last December and January. It was so encouraging to me that it seemed reasonable to try to mechanize the system immediately and use a computer to speed up further research on the linguistic side of the problem, rather than to perfect the linguistic system by hand, so to speak, and then mechanize it. The score still remains at 220 sentences. Since February, the computer programming needed to handle the essentially linguistic part of the system has been completed, and it is now in operation on ILLIAC. It does not look French words up in a mechanical dictionary. This seems to me to be an operation whose mechanization can legitimately be left until later. Quite closely similar problems must have been solved already for many other applications of computers. At any rate, I have to convert a French sentence by hand into a series of what would be the entries for the words in my mechanical dictionary. At the other end of the process, the computer produces a translation consisting of numbers that have to be looked up in a one-for-one table of English words. This, again, seems a legitimate and indeed trivial simplification in the development stage. My choice of French as the language to work on was obviously not dictated by a consideration of the market. The obvious choice is Russian; if I knew Russian, I too would have chosen it. However, my original intention was to demonstrate that a certain method of attack would yield results with surprising speed, and the method was one which second applicable to most pairs of languages. In common with most of those who have worked on mechanical translation, I have assumed that a set of rules could be devised by which most texts on a given subject in a given language could be transformed into texts in another language that were recognizably translations. One must also assume that the set of rules is not too large and complicated for human investigators to complete, or for a computer to apply. There are then two large questions; how to devise the rules, and how to enable the computer to apply them. In the beginning of my work I sat down to make up some rules before I had any clear idea of what sort of rules I would want; but in describing the approach now it is more convenient to begin by saying what sort of rules are involved. To begin with, it is assumed that the French words in a sentence have been looked up in a special dictionary, and that what I shall refer to as “items” have been brought out of the dictionary, one for each French word. Each item begins with a fixed number of digits that indicate the grammatical characteristics of the French word, in fairly conventional terms. Then comes a number indicating what the French word was, and then the English equivalent that will come out as part of the translation, unless it is changed in the course of working out the sentence. After that there may follow one or more instructions, then one or more constants that are used in carrying out instructions, and finally one or more diacritics whose presence or absence may be a necessary condition for executing various instructions. The sentence, in the form in which the computer gets it from the hypothetical dictionary search routine, contains instructions that will have to be carried out before translation is produced. These instructions correspond to the rules invented during the non-mechanical consideration of the problems. Some of the rules, however, are so general in their application that it is inefficient to plant them in individual dictionary items. For instance, there has to be a rule providing that adjectives, which mostly follow the nouns they modify in French, should be moved around to the English position, before the noun. This rule would apparently have to be included in the dictionary item for almost every French adjective. So it is more practical to have a number of general instructions, 12 of them at the moment, put at the beginning of each new sentence before instructions begin to be carried out. After each instruction is carried out, it is discarded; and when there are no instructions left, the English words remaining in the items of the sentence are printed out; they compose, one hopes, a translation of the original French sentence. An ordinary instruction is done once and then thrown away, but a general instruction has to be done once for each item in the sentence, it is treated as though it were found at the right moment in each item in turn. The order in which instructions are to be carried out has to be controlled very carefully, to avoid conflict. The most important reason for this is the need to make all the decisions that depend on French word order before the items are shuffled around into English word order. The sequence of execution is fixed by beginning each instruction with a priority number, within a somewhat arbitrary range of one to 126. Whenever the computer has to decide which instruction to follow next, it looks for the one with the lowest priority number. In case of a tie, the instruction occurring earlier in the sentence is done first. Within each item the instructions are listed in the order in which they are to be done, so that actually the computer only has to look at the first of the remaining instructions in each item, and at the first remaining one in the series of general instructions, to decide which one to take up next. Besides its priority number, seven binary digits in the present system, an instruction contains the name of an operation, nine digits, and a parameter, three digits. The parameter has a rather minor function, enabling reference to be made to comparison constants included in the same item with the instruction. To carry out the instruction, the name is used to refer to the operation, a series of consecutive computer words carried permanently in magnetic drum storage. The operation may begin with a series of one or more comparison constants which can be referred to by number, and after any such constants follows a series of words that might be called sub-instructions. These are converted by an interpretive routine into a program of sub-operations, which are computer routines permanently stored in the Williams memory. Such a program contains orders for making changes in the sentence when appropriate. It may contain logical decisions and loops of all kinds, but unlike a computer program it does not have much ability to alter itself. There are 48 basic sub-operations, and all the operations I have concocted or imagined so far can be conveniently programmed in terms of them by reference to simple tables, without having to do any computer programming. The various sub-operations can make the item containing the current instruction the “current” item, to be looked at or altered or moved, or can make the item before or after the presently current one into the new current one. They can ask whether the current item has the grammatical characteristics, or the English word, or the French word, indicated by one of the comparison constants; or whether it contains a given instruction or diacritic. Or they can look forward or backward in the sentence until they find an item that satisfies some such condition, making it the new current item if it is found, and returning a negative answer if none is found. Other sub-operations can change the grammatical characteristics or the English of an item, or insert an instruction or a diacritic into it, or delete or insert a whole item, or change the order of items in the sentence. The whole process of translation by this method can be described as a double interpretive routine. Raw material is got from the dictionary and then subjected to the first interpretive routine, which causes instructions to be performed in the correct order, and a series of English words to be printed out after the last instruction is performed. Each instruction, in turn, takes an operation from storage and interprets it as a program of sub-operations. Now the sub-operations, and the interpretive routines, are so general as to be almost independent of what languages are concerned in the translation. So the method has the possible advantage that one master program, with only slight changes, might be used for several different sorts of translation. A different dictionary would have to be used in each case, of course, and a different set of instructions would have to be stored on the drum. But, as the basic program has taken several months of part time work to get ready, it is encouraging to think that this work may not have to be repeated if I should get mechanical translation of chemical French into actual operation, and then turn to some language more in demand at the present time. It remains to describe the method by which the rules, and the dictionary items for them to work on, have been arrived at. I opened a recent French chemical journal at random, went to the beginning of the article, and set out to formulate verbal rules that would translate the first sentence. It had about forty words, and it took ten hours to work out the rules. Turning to the second sentence, I added new items to the dictionary invented new rules, and

Read the paper · More papers on PaperTik