Multi-class Document Classification Using Improved Word Embeddings

Benedict A. Rabut, Arnel C. Fajardo, Ruji P. Medina · 2019

In this paper, we conducted an experiment to build a classification model that combines different techniques in most of the Natural Language Processing Tasks. We used the word embedding method to transform every word in the dataset and to obtain the custom-built word embedding vectors. This is in contrast to the approaches in the previous literature that implement word embedding using the pre-trained word embedding vectors. We enriched the custom-built word embedding vectors by incorporating Part-of-Speech (POS) tag vectors to provide additional semantic information about the word to be used in training our proposed classification model. The proposed model was built using the neural network approach, which is considered to be more efficient and reliable in solving real problems for document classification tasks. We fine-tuned the parameters during the training of our neural network classification model with our aim to increase the performance in terms of classification accuracy. The experimental result demonstrates that our model performs remarkably well and increase the percentage accuracy up to 1.7% compared to the accuracy results obtained by the previous baseline word embedding methods using the same dataset. It was also observed that our model outperforms some other traditional classification models implemented using different techniques and machine learning algorithms.

Read the paper · More papers on PaperTik