An Augmented Translation Technique for Low Resource Language Pair

Rashi Kumar, Piyush Jha, Vineet Sahula · 2019

Neural Machine Translation (NMT) is an ongoing technique for Machine Translation (MT) using an enormous artificial neural network. It has exhibited promising outcomes and has shown incredible potential in solving challenging machine translation exercises. One such exercise is the best approach to furnish great MT to language sets with a little preparing information. In this work, we inspect Zero-Shot Translation (ZST) for a low resource language pair. By working on high resource language pairs for which benchmarks are available, namely Spanish to Portuguese, and training on data sets(Spanish-English and English-Portuguese), we prepare a state of proof for ZST system that gives appropriate results on the available data. Subsequently, we test the same architecture for Sanskrit to Hindi translation for which data is sparse, by training the model on English-Hindi and Sanskrit-English language pairs. To prepare and decipher with the ZST system, we broaden the preparation and interpretation pipelines of the NMT seq2seq model in TensorFlow, incorporating ZST features. Dimensionality reduction of word embedding is performed to reduce the memory usage for data storage and to achieve faster training and translation cycles. In this work, existing helpful technology has been utilized imaginatively to execute our Natural Language Processing (NLP) issue of Sanskrit to Hindi translation. We have constructed a Sanskrit-Hindi parallel corpus of 300 sentences for testing. The data required for the construction of a parallel corpus has been taken from the telecasted news, published on the Department of Public Information, the state government of Madhya Pradesh, India website.

Read the paper · More papers on PaperTik