Normalising the IJS-ELAN Slovene-English Parallel Corpus for the Extraction of Multilingual Terminology.

Gaël Dias, Špela Vintar, Gabriel Pereira Lopes, Sylvie Guilloré · 1999

Various efforts have been made for the development of tools and methods dedicated to the automatic processing of multilingual terminology databases. For that purpose, multilingual parallel corpora have been used as a basis resource. However, most of the neologisms in technical and scientific domains are realised by multiword terms that are rarely identified in parallel corpora. In this paper, we propose the normalisation of the IJS-ELAN Slovene-English parallel corpus by using the language-independent SENTA software. 1 Introduction The need for multilingual terminology resources has become particularly acute owing to the globalization of scientific and technical exchanges and the concurrent development of international communication networks. As a consequence, various efforts have been made for the development of tools and methods dedicated to the automatic processing of multilingual terminology databases. For that purpose, multilingual parallel corpora have been used as a...

Read the paper · More papers on PaperTik