Part-of-speech Tagging of Code-Mixed Social Media Text
Souvick Ghosh, Satanu Ghosh, Dipankar Das · 2016
A common step in the processing of any text is the part-of-speech tagging of the input text.In this paper, we present an approach to tackle code-mixed text from three different languages Bengali, Hindi, and Tamilapart from English.Our system uses Conditional Random Field, a sequence learning method, which is useful to capture patterns of sequences containing code switching to tag each word with accurate part-of-speech information.We have used various pre-processing and post-processing modules to improve the performance of our system.The results were satisfactory, with a highest of 75.22% accuracy in Bengali-English mixed data.The methodology that we employed in the task can be used for any resource poor language.We adapted standard learning approaches that work well with scarce data.We have also ensured that the system is portable to different platforms and languages and can be deployed for real-time analysis.