Cross-lingual sentence extraction for information distillation

Adish Kumar Singla, Dilek Zeynep Hakkani-Tür · 2008

Information distillation aims to analyze and interpret large vol-umes of speech and text archives in multiple languages and pro-duce structured information of interest to the user. In this work, we investigate cross-lingual information distillation, where non-English (source language) documents are searched for user queries that are in English (target language). We propose to per-form distillation both on the original source language data and their English translations output by machine translation, and combine the two outputs. We experimentally show that com-bination approach results in 8 % to 16 % absolute (13 % to 31% relative) F-measure improvement over the previous work. Index Terms: information distillation, sentence extraction, cross-lingual processing, and classification model combination.

Read the paper · More papers on PaperTik