Zero-shot approach for news and scholarly article classification
Devika Patadia, Shivam Kejriwal, Pashva Mehta, Abhijit R. Joshi · 2021
With the enormous growth of data and documents in raw and unstructured formats, categorizing data is the need of the hour. Understanding text is an essential step in classifying it. Natural Language Processing (NLP) aids computers in comprehending and interpreting natural languages such as English, allowing them to evaluate and deduce their meaning. Keeping in mind large, diverse data at ever-increasing rates, this paper compares Zero-shot learning, a relatively new technique for classifying text, to multinomial Naive Bayes, which is a traditional classification method. The experiment analysis has been carried out on the BBC News dataset and the ArXiv dataset, which contains metadata of scholarly articles. The Zero-shot approach gives promising results with top-2 accuracy of about 93% and 87.6% on the two unseen datasets and can hence be used to effectively classify unlabelled data without specific training.