Word embedding with hellinger PCA to detect the sentiment of bengali text
Md. Saiful Islam, Md. Al-Amin, Shapan Das Uzzal · 2016
The sentiment of a sentence or a comment can be detected more accurately by applying Word Embeddings. This article presents the idea of word co-occurrence matrix and Skip-Gram to determine the actual contexts of the words, Hellinger PCA to determine the most similar words and generate a sliding window of most probable context words around each word. It is shown that, by applying Word Embeddings to classify the sentiment of a comment achieves higher accuracy with larger corpus. For our corpus of 2500 comments, the accuracy achieved is 70%, which is rapidly increasing with the size of the corpus.