Deep Learning methods for Subject Text Classification of Articles

Piotr Semberecki, Henryk Maciejewski · Annals of Computer Science and Information Systems · 2017

This work presents a method of classification of text documents using deep neural network with LSTM (long shortterm memory) units.We have tested different approaches to build feature vectors, which represent documents to be classified: we used feature vectors constructed as sequences of words included in the documents, or, alternatively, we first converted words into vector representations using word2vec tool and used sequences of these vector representations as features of documents.We evaluated feasibility of this approach for the task of subject classification of documents using a collection of Wikipedia articles representing 7 subject categories.Our experiments show that the approach based on an LSTM network with documents represented as sequences of words coded into word2vec vectors outperformed a standard, bag-of-word approach with documents represented as frequency-of-words feature vectors.

Read the paper · More papers on PaperTik