Building and processing a multilingual corpus of parallel texts

Peter Stahl · 2002

The minimum requirements in building a multilingual corpus of parallel texts are tags in the texts that establish cross-references between corresponding parts. These tags can be used to merge several independent text files into one, displaying aligned units of text, in order to do research with the help of the TUSTEP word processor. Extensive use is made of the complex possibilities of pattern matching: words, word forms, prefixes, suffixes, and one or more explicit or abstract character strings can be searched, while at the same time excluding others, by entering instructions into the word processor interactively. By doing so, the user does not depend upon predefined tags which contain semantic or grammatical information. Examples taken from the Finnish-German parallel corpus show how such instructions are written and what the results look like. Apart from interactive work, pattern matching and text processing tasks can be done by executing parameter driven program files. Examples show how four text files are aligned to produce a POSTSCRIPT as well as an HTML file.

Read the paper · More papers on PaperTik