Reducing the Impact of Data Sparsity in Statistical Machine Translation

Karan Singla, Kunal Sachdeva, Srinivas Bangalore, Dipti Misra Sharma, Diksha Yadav · 2014

Morphologically rich languages generally require large amounts of parallel data to adequately estimate parameters in a statistical Machine Translation(SMT) system.However, it is time consuming and expensive to create large collections of parallel data.In this paper, we explore two strategies for circumventing sparsity caused by lack of large parallel corpora.First, we explore the use of distributed representations in an Recurrent Neural Network based language model with different morphological features and second, we explore the use of lexical resources such as WordNet to overcome sparsity of content words.

Read the paper · More papers on PaperTik