Hindi Text Classification: A Review

Manan Mayur Mehta, Utkarsh Pandey, Yash Chaudhary, Raghav Sharan Sharma, Ishaan Gill, Deepak Kumar Gupta, Ashish K. Khanna · 2021 3rd International Conference on Advances in Computing, Communication Control and Networking (ICAC3N) · 2021

In recent years, Natural Language Processing (NLP) and text analysis has been a topic of great interest among researchers. The utilization of deep learning techniques such as LSTM, Bi-LSTM, CNN and Transformers in text handling has given us remarkable results. In this work, we review a large group of deep learning designs for text classification tasks, particularly for Hindi text. Due to the absence of a large labelled corpus of Hindi language, the research has been limited. IndicFT FastText embeddings were used in this work as they have surpassed available pre-trained embeddings on several machine learning tasks. Transformer models employed in this work are mBERT, IndicBERT and XLM-RoBERTa.

Read the paper · More papers on PaperTik