Character Level Neural Architectures for Boosting Named Entity Recognition In Code Mixed Tweets
Abhishek Narayanan, Abijna Rao, Abhishek Prasad, Bhaskarjyoti Das · 2020 International Conference on Emerging Trends in Information Technology and Engineering (ic-ETITE) · 2020
Named Entity Recognition (NER) is a highly researched topic in the field of Natural Language Processing and is an important step towards information extraction. The rampantly growing popularity of Online Social Networks (OSNs) has led to an exponential increase in the amount of raw text comments online that provides a significantly fertile ground to analyze data and procure certain insights. However, most of the state-of-the-art models existing at present which have achieved near human performance, are primarily targeted at or cater to the mining and analysis of resource rich monolingual text data. Code-mixing or multilingualism in OSNs results in significantly unstructured text which poses a serious problem for social media analytics by hindering algorithms from learning to generalize on the unstructured data. With the motive of tackling this problem and boosting the performance of NER on code mixed text, this paper proposes the usage of word embeddings and specifically character level recurrent neural networks to take into account the context in which words are used. This approach facilitates effective segmentation and classification of text into entities like person, location or organization for text data composed of Hindi-English Code mixed tweets. Our extensive experimentation indicates an improvement over the minimal existing baseline state-of-art approaches for extracting named entities.