Topic categorization of Tamil News Articles using PreTrained Word2Vec Embeddings with Convolutional Neural Network

S. Ramraj, R. Arthi, Solai Murugan, Montee Ste Julie · 2020

Almost all the problems in NLP are solved using various techniques from machine learning to Deep Learning. Still, there is mystery in language localization. NLP problems are unclear for languages other than English. The problems may be named as Entity Extraction, OCR or classification and prediction in sequence modelling. The amount of people using local language (Tamil, Telegu, Hindi etc) in the social media is increasing, so it is important to automate the process of classifying those contents. Here, the aim is to classify the Tamil news articles to its related topics (Sports, Cinema, Politics). In the existing work they have approached traditional machine learning methods with TFIDF of words as features. In this work we have compared the existing TFIDF feature learning along with Pre-Trained embeddings given to Convolutional Neural Networks (CNN). We found that CNN with pretrained embeddings gave better F1 score compare to TFIDF feature learned with Support Vector Machine (SVM), Naive Bayes (NB) algorithm.

Read the paper · More papers on PaperTik