Dialect Recognition in Tamil and Telugu: An Integrated Approach

G Naveen Raaghavendran, P Jaswanth, S Shreenithi, Sri Vaishnavi J V, K. R. Bindu · 2025

Dialect recognition is one of the important areas in NLP, which focuses on identifying regional differences within a language. This work explores various approaches, ranging from traditional machine learning algorithms to deep learning architectures, large language models, and insights from previous research to develop high-accuracy dialect recognition systems respectively. Whisper model performed the best with an accuracy of 93.18% in dialect recognition was achieved, demonstrating its effectiveness in capturing linguistic variations. HuBERT achieved 92.86% accuracy but struggles with classification between similar dialects. The aim is to improve dialect recognition accuracy and enhance language understanding through the analysis of speech patterns, acoustic features, and linguistic characteristics. This research responds to the critical need for dialect-specific language resources, bridging gaps in speech recognition and machine translation. Our work makes an important contribution based on better communication and social cohesion among speakers of the Tamil and Telugu languages.

Read the paper · More papers on PaperTik