Using word signature features for terminology translation from large corpora

Pascale Fung · 1997

Automated translations of technical and domain-specific terms is highly desirable for machine translation systems since such terms are usually not found in standard dictionaries nor are they easily translatable by humans. This thesis describes a statistical approach, augmented with linguistic knowledge, to terminology translation. Most other statistical work on terminology translation uses word frequencies and occurrence positions in sentence-aligned translated texts as translation features. However, sentence-to-sentence translated texts are not very common. More irregular (noisy) translated texts and monolingual texts are easier to find. Unfortunately, word frequency and occurrence features are susceptible to translation or OCR noise, and are not applicable to monolingual texts. The focus of this thesis is to extract discriminatory statistical word signature features applicable to noisy parallel texts and same domain, nonparallel texts. A translator aid tool is also proposed for terminology translation using noisy parallel corpora. We also explore the new possibility of using same domain, monolingual nonparallel texts of a pair of languages for extracting translation. We present two major types of signature features representative of domain-specific terminology from corpora: (1) the nonlinear K-vec and dynamic K-vec (DK-vec) features for noisy parallel corpora, and (2) the Word Relation Matrix (WoRM) feature for nonparallel corpora. We also present various matching techniques and similarity measures for finding associated bilingual word pairs. Using the first type of feature on noisy parallel corpora, our evaluations show a 55.35% precision from a small corpus and 89.93% precision from a larger corpus. Humans achieved a 47% increase in accuracy using the system output, when translating domain-specific terms. The evaluation results of using the Word Relation Matrix alone show a precision of about 30%. We show that this is a useful result since human translation accuracy is increased by about 51%. We believe our results indicate a promising direction for exploring the vast quantity of online nonparallel material for terminology translation. We also discuss possible extensions to improve these results and potential applications of our features in other areas of natural language processing. As terminology extraction is a preprocessing step to translation, we also present two types extraction algorithms that we developed. Due to some important differences between the Chinese and Japanese, we adopt totally different approaches for the two languages--lexical approach for Chinese and a grammatical approach for Japanese.

Read the paper · More papers on PaperTik