Recognizing non-translatable symbols in a multi-lingual computer--assisted translation system for DTP documents

Szymon Grabowski, Cezary Draus, Wojciech Bieniecki · Automatyka / Akademia Górniczo-Hutnicza im. Stanisława Staszica w Krakowie · 2010

Such programs work on small text segments (snippets) like sentences or expressions. Translated segments are stored in a database (called Translation Memory, TM [1]), ready to be used again. If a segment to be translated is found in the existing database (exact matching), the corresponding translation is proposed to be reused. If more than one translation of a given snippet is found, then several suggestions are proposed and it is up to the translator to decide which one to apply. Some CAT tools work with raw files, while others are designed to handle more sophisticated formats of documents like word-processing, desktop publishing (DTP) files, web pages or presentation documents. Among them, desktop publishing documents represent a great challenge to the CAT tool family, because of their graphical diversity. A 2006 survey [2] among 874 translation professionals from 54 countries revealed that although over 80% of the polled individuals do work with a translation memory system, most of them do not use a TM tool for all content, and the main reported reasons were: hardcopy documents available only; lack of support for the desired file format; the TM tools are too complicated for short texts; the repetition rate is too low; the tool is not suitable for the user’s text types; complex layout. Obviously, better and more user-friendly applications could increase the use of this kind of software.

Read the paper · More papers on PaperTik