Aligning verb + noun collocations to improve a French-Romanian FSMT system

Amalia Todiraşcu, Mirabela Navlea · Amsterdam studies in the theory and history of linguistic science. Series 4, Current issues in linguistic theory · 2018

Abstract We present several Verb + Noun collocation integration methods using linguistic information, aiming to improve the results of a French-Romanian factored statistical machine translation system (FSMT). The system uses lemmatised, tagged and sentence-aligned legal parallel corpora. Verb + Noun collocations are frequent word associations, sometimes discontinuous, related by syntactic links and with non-compositional sense ( Gledhill, 2007 ). Our first strategy extracts collocations from monolingual corpora, using a hybrid method which combines morphosyntactic properties and frequency criteria. The second method applies a bilingual collocation dictionary to identify collocations. Both methods transform collocations into single tokens before alignment. The third method applies a specific alignment algorithm for collocations. We evaluate the influence of these collocation alignment methods on the results of the lexical alignment and of the FSMT system.

Read the paper · More papers on PaperTik