Automatic Dictionary Construction and Identification of Parallel Text Pairs
Sumithra U. Velupillai, Martin Hassel, Hercules Dalianis · 2008
When creating dictionaries for use in for example cross-language search engines, parallel or comparable text pairs are needed. Multilingual web sites may contain parallel texts but these can be difficult to detect. For instance, a multilingual website, Hallå Norden, contains information in five languages; Swedish, Danish, Norwegian, Icelandic and Finnish. Working with these texts we discovered two main problems: the parallel corpus was very sparse, containing on average less than 80.000 words per language pair (in the final version of the corpora), and it was difficult to automatically detect parallel text pairs. We discovered that, on average, around 55 percent of the texts were not parallel. Creating dictionaries with the word aligner Uplug gave on average 213 dictionary entries. Despite the corpus sparseness the results were surprisingly good compared to other experiments with larger corpora. Following this work, we made two sets of experiments on automatic identification of parallel text pairs. The first experiment utilized the frequency distribution of word initial letters in order to map a text in one language to a corresponding text in another in the JRC-Acquis corpus (European Council legal texts). Using English and Swedish as language pair,