A survey on Gujarati NLP research work
Brijeshkumar Y. Panchal, Apurva Shah · Aibi revista de investigación administración e ingeniería · 2025
Natural Language Processing (NLP) is an area of Artificial Intelligence (AI) that agreements with text, speech, and translation. Internet users of vernacular languages have been dramatically increasing day by day. Therefore, NLP researchers have been working in regional languages; and, Gujarati is one of them. Gujarati is an Indo-Aryan language inborn to the Indian state of Gujarat and vocalized by Gujarati individuals. Around 62 million Gujarati speakers are all around the world that ranks as the 26th most extensively spoken language globally. However, Gujarati is the youngest and lowest-resource Indian language in NLP community. However, little path-breaking work has been done in Gujarati NLP (GNLP). E.g., WordNet, Morphological, Stemmer, optical character recognition (OCR), Speech Recognition, Parts of Speech, Machine Translation, etc. Many researchers have been working with a rule-based approach for GNLP. After that, only a few researchers have attempted the machine learning, deep learning, and reinforcement learning approaches. This paper focuses on a critical survey of existing GNLP research, covering research papers from 1999 to August 2024 in this study. This survey predicts GNLP study until 2030 with the use of available data. This Study explores gaps in present studies, and suggestions for the newly, active research field of GNLP, through this paper one can decided to develop Deep Learning based GNLP system, to get more accuracy.