Evaluation and Comparison of Cross-lingual Text Processing Pipelines

Robert Jungnickel, André Pomp, Andreas Kirmse, Xiang Li, Владимир Самсонов, Tobias Meisen · 2019

With the trend of globalization and digitalization, many transnational companies are continuously collecting and storing unstructured text data in different languages. To exploit the business value of such high-volume multilingual text data, cross-lingual information extraction utilizes machine translation and other natural language processing (NLP) techniques to analyze this data. However, results of these analysis heavily depend on the order in which the tasks are performed as well as the used machine translation and NLP approaches or trained models. In this paper, we defined and evaluated a series of cross-lingual text processing pipelines for English and Chinese language. We therefore combine multiple commercial machine translation services with different automatic keyphrase extraction and named entity recognition techniques and evaluate their performance with regards to the order of execution. Hence, we evaluate the combination of machine translation systems and natural language processing techniques with two processing sequences in our experiment. One is to translate the document before extracting keyphrase and named entities. The other is to translate the processing results. The experiment outcomes indicate that translating documents is a better choice than the other way around in both tasks. However, there exists a substantial disparity between the performance of the cross-lingual text processing pipelines and the corresponding monolingual references.

Read the paper · More papers on PaperTik