Evaluating Machine Learning and Deep Learning Techniques for Part-Of-Speech Tagging in Tamil

Kiruthika Raja, Pawan Kumar Chaurasia, Midhunchakkaravarthy · International Journal of Basic and Applied Sciences · 2025

Part-of-speech (POS) tagging for Tamil is a importance task due to the language’s highly inflectional and agglutinative morphology. This ‎study systematically evaluates both machine learning and deep learning models including Conditional Random Fields (CRF), Support Vector Machine (SVM), Hidden Markov Model (HMM), Long Short-Term Memory Recurrent Neural Network (LSTM-RNN), and LSTM-RNN with CRF output for Tamil POS tagging, using a well-annotated CLE-style benchmark dataset. We employed a comprehensive, lan-‎language-independent feature set and performed 10-fold cross-validation to ensure robust results. Experimental finding that, for the CLE da-‎dataset, the CRF model achieves the highest average accuracy at 86.32%, outperforming SVM (81.13%), LSTM-RNN (78.64%), LSTM-RNN-CRF (78.03%), and HMM (78.03%). In contrast, on the more challenging BJ dataset, the LSTM-RNN deep learning model attains ‎the highest accuracy of 92.70%, followed closely by CRF (91.2%), LSTM-RNN-CRF (91.02%), HMM (90.11%), and SVM (86.25%). ‎These results highlight the importance of model selection in morphologically rich languages, while CRF is optimal for structured and moderately sized datasets, LSTM-RNN deep learning approaches excel on larger. This work establishes new empirical benchmarks for Tamil ‎POS tagging and demonstrates that advanced neural models provide a clear advantage in handling Tamil’s linguistic complexity‎.

Read the paper · More papers on PaperTik