Handling terminology-intensive texts in machine translation
Eleni Efthimiou, Marianna Katsoyannou · RIAO Conference · 2000
When post-editing the MT output of terminology-intensive documents, three problem sources were identified: i) Their terminology content ii) The structure of multi-word terms inside and across languages or terminology domains iii) The one-to-many translations of some terms To resolve the term boundary problem, we first tried an on-processing recognition based on syntactic analysis and the [+TERM] value of the lexical head. Although effective, the full scale application of this solution adds a considerable load to a heavy grammar, while the problem of identifying the boundary of not yet fixed terms remains unsolved. Alternatively, when a number of options are available in respect to recognizing parts of a linguistic string in the context of a term head as different terms, the system asks the user to mark his/her preference. The marked string is checked against the term DB and processing continues unambiguously. If no correspondence in the term DB is found the selected sting receives an index which identifies it as a term, a property curried to the translation object. To address the problem of one-to-many term translations/correspondences, we incorporated the option of user interaction in pre-processing for the case of single-word terms, so that selected correspondences be marked and saved for use before the actual linguistic analysis.