Evaluating Machine Learning and Deep Learning Techniques for Part-Of-Speech Tagging in Tamil
Kiruthika Raja, Pawan Kumar Chaurasia, Midhunchakkaravarthy · International Journal of Basic and Applied Sciences · 2025
Part-of-speech (POS) tagging for Tamil is a importance task due to the language’s highly inflectional and agglutinative morphology. This study systematically evaluates both machine learning and deep learning models including Conditional Random Fields (CRF), Support Vector Machine (SVM), Hidden Markov Model (HMM), Long Short-Term Memory Recurrent Neural Network (LSTM-RNN), and LSTM-RNN with CRF output for Tamil POS tagging, using a well-annotated CLE-style benchmark dataset. We employed a comprehensive, lan-language-independent feature set and performed 10-fold cross-validation to ensure robust results. Experimental finding that, for the CLE da-dataset, the CRF model achieves the highest average accuracy at 86.32%, outperforming SVM (81.13%), LSTM-RNN (78.64%), LSTM-RNN-CRF (78.03%), and HMM (78.03%). In contrast, on the more challenging BJ dataset, the LSTM-RNN deep learning model attains the highest accuracy of 92.70%, followed closely by CRF (91.2%), LSTM-RNN-CRF (91.02%), HMM (90.11%), and SVM (86.25%). These results highlight the importance of model selection in morphologically rich languages, while CRF is optimal for structured and moderately sized datasets, LSTM-RNN deep learning approaches excel on larger. This work establishes new empirical benchmarks for Tamil POS tagging and demonstrates that advanced neural models provide a clear advantage in handling Tamil’s linguistic complexity.