Leveraging Comparable Corpora for Computer‐assisted Translation

Estelle Maryline Delpech · 2014

This chapter starts with a historical approach to computer-assisted translation (CAT). It retraces the beginnings of machine translation and explains how the CAT has developed so far, with the recent appearance of the issue of comparable-corpus leveraging. The chapter explains the current techniques to extract bilingual lexicons from comparable corpora. It provides an overview of the typical performances, and discusses the limitations of these techniques. The chapter describes the prototyping of the CAT tool meant for the comparable corpora and based on these techniques. Specific approaches have been developed to acquire bilingual lexicons from comparable corpora. There are methods based on frequency distribution or the use of semantic relations. The CAT prototype is able to extract terms from texts in the source and target languages and align them with a method based on a distributional approach.

Read the paper · More papers on PaperTik