A Generic Approach to Processing Parallel Corpora of the Europarl for Distributional Discourse Patterns

Jolanta Mizera–Pietraszko · International Journal of Intelligent Computing Research · 2011

This work presents a generic approach to building a featured test set in the light of further query processing by multilingual information retrieval systems.Language pair phenomena are found crucial points both in the process of query translation and document retrieval.Therefore, our approach aimed at improvement of machine translation quality is presented on the base of the French aligned to English version of the Europarl corpora.We investigate frequency of the language discourse occurrences commonly used in speech and writing.Linguistic structures are extracted from the corpora together with their representatives in order to create a featured test set of English and French grammatical patterns.In multilingual retrieval process a selection of patterns with the highest frequency simultaneously in these two languages constitutes an indication to a user of such a query formulation that may result in the most relevant system responses.This study of linguistic properties shows how to make the query easier for automatic translation and consequently to improve the system responsiveness.

Read the paper · More papers on PaperTik