Bilingual Parallel Corpora: A Major Resource for Developing Computational Tools for Automatic Processing of Hindi-Dogri Language Pair

Joginder Kumar, Manik Rakhra, Preeti Dubey · 2022 10th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO) · 2022

Dogri is a language with low computational resources. There are hardly any digital resources available for this language. The major hindrance to developing these resources is the lack of corpora. This paper describes the authors' attempt to create a bilingual parallel corpus of 0.2 million Hindi to Dogri sentences to develop various natural language processing tools like statistical & neural machine translation systems, summarization, classification, romanization, etc. These tools require a sufficient amount of corpus as a dataset when developed using deep neural models. The corpus was created for the Hindi-Dogri language pair, with a significant emphasis on ambiguous words to resolve ambiguity using deep learning models. The paper also highlights the limitations of the available rule-based Hindi to Dogri machine translation system with respect to wrong translation results and ambiguity in the output, emphasizing the need to develop systems by training on a corpus with ambiguous data using deep learning techniques.

Read the paper · More papers on PaperTik