Lost in specialised translation: the corpus as an inexpensive and under-exploited aid for language service providers
Gloria Corpas Pastor · 2007
Most translation activities nowadays involve the rendering of non-literary texts in the context of (semi-)specialised communication. The vast majority of translated texts usually belong to rather specialised and repetitive genres, which makes them particularly suitable for multilingual document management and CAT tools. In fact, translation memory systems (TMS) have proven an indispensable tool for professional translators in that they have been able to enhance the efficiency and cost-effectiveness of translation of voluminous repetitive texts without compromise of quality. Whereas for someone who works on translation of highly repetitive texts from a narrow domain TMS provide a smooth and problem-free solution, translators working on a vast range of different specialised texts may find that things are not as easy as they seem. First, parallel corpora (TM files) may be either nonexistent or difficult to obtain. Secondly, the growth of TMS cannot catch up with the growth of bilingual corpora or bilingual websites. Neither can they catch up with being representative of dynamically developing domains where new terminology is being proposed on a daily basis. This is where other freely available — or relatively inexpensive — resources could come into play. In this paper we will focus on the compilation of virtual or ad-hoc corpora (i.e. corpora mined from electronic sources for a specific task) and the tools enabling their exploitation by translators. We will demonstrate how a step-by-step approach to building an adequate and/or representative corpus from resources in the Internet works in practice. Corpus design criteria and qualitative issues will be taken into account. For illustrative purposes, we will discuss a translation assignment of a specialised nature. Real examples will be presented that show how to mine a corpus and how to use it in order to meet translators’ needs as far as terminology, documentation, target text conventions and other constraints are concerned. We will illustrate how language service providers can use freeware concordances to search for translation equivalents in comparable and parallel corpora.