Data Pre-Processing Framework for Kannada Vachana Sahitya
C. B. Lavanya, H.S. Nagendra Swamy, Pavan Kumar M.P. · 2024
Advancements in Natural Language Processing (NLP) driven by machine learning, deep learning, and artificial intelligence have significantly broadened its scope and improved interactions between humans and computers. Despite these advancements, NLP systems encounter challenges arising from incomplete and error-prone data, which can result in biased model outputs. Technical domains present further hurdles, necessitating domain-specific fine-tuning and the development of custom lexicons. Additionally, many languages lack robust NLP support, limiting accessibility. In this context, innovative NLP data pre-processing and tokenization methods tailored for Kannada Vachana Sahitya texts are investigated. The majority of documents or text utilized in any language processing applications consists of raw text, in which some of the words are not represented in the standard form. There is a necessity of pre-processing the input text before building the translation model. In this paper, the main goal is to pre-process the source text with essential steps like Cleaning, Parsing, Tokenization, Padding, Stemming and Lemmatization. During this process, Kannada source text - Vachana Sahitya is tested on bilingual corpus and approximately 35% of non-standard input text is considered for experimental analysis.