Exploiting alignment in multiparallel corpora for applications in linguistics and language learning

Johannes Graën · Zurich Open Repository and Archive (University of Zurich) · 2018

This thesis exploits the automatic identification of semantically corresponding units in parallel and multiparallel corpora, which is referred to as alignment.Multiparallel corpora are text collections of more than two languages that comprise reciprocal translations.The contributions of this thesis are threefold:• First, we prepare a large multiparallel corpus by adding several layers of annotation and alignment.Annotation is first performed on each language individually, while alignment is applied to two or more languages.For the latter case, we use the term multilingual alignment.We show that word alignment on parallel corpora can improve language-specific annotation by means of disambiguation.• Our second contribution consists in the development and evaluation of prototypical algorithms for multilingual alignment on both sentence and word level.As languages vary considerably with regard to how content is realized in sentences and words, multilingual alignment needs to be represented by a hierarchical structure rather than by bidirectional links as prevailing representation of bilingual alignment.• Based on our corpus, we thirdly show how word alignment in combination with different types of annotation can be employed to benefit linguists and language learners, among others.All tools developed in the context of this thesis, in particular the publicly available web applications, are driven by efficient database queries on a complex data structure.iii First of all, I wish to thank my supervisor Martin Volk who guided me through the initial troubles, gave me room to realize my own ideas and had the necessary confidence in me to finish this big project of mine.I am likewise

Read the paper · More papers on PaperTik