Natural Language Processing based Text Imputation for Malayalam Corpora

Annlin Rojan, Edwin Alias, Georgy M. Rajan, Jithin Mathew, Dhanya Sudarsan · 2020 International Conference on Electronics and Sustainable Communication Systems (ICESC) · 2020

The prediction task in Natural Language Processing intends to figure out the missing characters, letter, word, expression, or sentence that conceivable follows in a given fragment of a book. Since the beginning of NLP, numerous frameworks with various techniques were produced for various dialects. The missing content prediction is one among the significant concerns of Natural Language Processing. Moreover, most of the text prediction related tasks are conducted in different dialects but not in Malayalam. Over these years though having some of the best classics, historical records and many more to keep the world interested with, the Malayalam was lacking many of the advantages that any other language process in the digital world. This is because there are only a few standard models that could fit into the Malayalam. In this paper, an attempt is made to fit the BERT into Malayalam. A Malayalam pre-trained model will be creating and implementing some of its applications with the model and discussing its scope and areas of development to be addressed and future research scope.

Read the paper · More papers on PaperTik