Advancing NLP for Punjabi Language: A Comprehensive Review of Language Processing Challenges and Opportunities
Gurpej Singh, Rahul Bhandari, Prabhdeep Singh · 2024
Punjabi, an Indo-Aryan language spoken by mil-lions, remains significantly underrepresented in the field of natural language processing (NLP). Despite its extensive speaker base, Punjabi has not received equivalent NLP research and development attention compared to other major languages. This review paper delves into the existing gaps in Punjabi NLP research, identifying key challenges such as the acute scarcity of digital resources, unique complexities due to the language's capi-talization patterns, and the absence of a standardized stopwords list. It also highlights the need for development in advanced areas like robust Part-of-Speech (POS) tagging, effective tokenization methods, and the complexities of managing an agglutinating language. Furthermore, we underscore the challenges in stem-ming and lemmatization processes specific to Punjabi and the critical need for labeled datasets for tasks like sentiment analysis. The primary objective of this review is to shed light on these challenges, thereby paving the way for future research initiatives that aim to enhance Punjabi's representation and application in NLP advancements.