Generation of noun-noun compounds in the Spanish-English machine translation system SPANAM®

Julia Aymerich · 2001

The translation of Spanish Noun + preposition + Noun (NPN) constructions into English Noun-Noun (NN) compounds in many cases produces output with a higher level of fluency than if the NPN ordering is preserved. However, overgeneration of NN compounds can be dangerous because it may introduce ambiguity in the translation. This paper presents the strategy implemented in SPANAM to address this issue. The strategy involves dictionary coding of key words and expressions that allow or prohibit NN formation as well as an algorithm that generates NN compounds automatically when no dictionary coding is present. Certain conditions specified in the algorithm may also override the dictionary coding. The strategy makes use of syntactic and lexical information. No semantic coding is required. The last step in the strategy involves post-editing macros that allow the posteditor to quickly create or undo NN compounds if SPANAM did not generate the desired result. Keywords Noun-Noun compounds, Spanish-English machine translation and SPANAM system SPANAM today SPANAM has been operational at the Pan American Health Organization since 1980. The program, along with its English-Spanish counterpart Engspan, was ported from the mainframe to the PC in 1992 and then to the Windows environment in 2000 (León, 2000). The basic architecture of the program hasn't changed since 1985: it is a transfer system with an ATN that generates a top-down, left-to-right sequential parse with chronological and explicit backtracking (Vasconcellos and León, 1988; Amores, 1996). Dictionary entries are rich in morphological, syntactic, and semantic information. The grammar and dictionaries have grown considerably in the past 15 years. SPANAM dictionaries currently contain over 95,000 source entries and 85,000 translations. Source entries include single words (65,000 stem forms), multiple word entries or SUs1 (11,000), analysis rules or AUs (8,500), and context-sensitive translation rules (10,000). In several MT evaluation experiments sponsored by the

Read the paper · More papers on PaperTik