Using topic models in domain adaptation

Samira Tofighi Zahabi, Somayeh Bakhshaei, Shahram Khadivi · 2014

An important factor of a corpus is its domain, usually the quality of a SMT system trained on an in-domain corpus increases by adding out-of-domain sentences to its training corpus. In this paper we have shown out-of-domain corpora may also contains sentences which are proper for improving the quality of in-domain corpus. These sentences have words and phrases that occur in indomain corpora so, their context is more similar to the context of in-domain parallel corpus and is far from context of out-of-domain parallel corpora. In this paper we suggest a method based on topic models to extract some sentences from out-of-domain parallel corpora that their context are similar to indomain parallel corpus. We used these extracted sentences for training an SMT system. Finally, we will show the BLEU score of the system output increases about 4.69% by adding these extra information to its training corpus.

Read the paper · More papers on PaperTik