A Novel Pipeline for Domain Detection and Selecting In-domain Sentences in Machine Translation Systems
Javad Pourmostafa Roshan Sharami, Dimitar Shterionov, Pieter H.M. Spronck · Research portal (Tilburg University) · 2021
General-domain corpora are becoming increasingly available for Machine Translation (MT) systems. However, using those that cover the same or comparable domains allow achieving high translation quality of domain-specific MT. It is often the case that domain-specific corpora are scarce and cannot be used in isolation to effectively train (domain-specific) MT systems. This work aims to improve in-domain MT by (i) a novel unsupervised pipeline for identifying distributions of different domains within a corpus and (ii) a data selection technique that leverages in-domain monolingual or parallel data to select domain-specific sentences from general corpora according to the distribution defined in (i).