Improving Short Text Classification Using Fast Semantic Expansion on Multichannel Convolutional Neural Network
Natthapat Sotthisopha, Peerapon Vateekul · 2018
Nowadays, text classification is recognized as one of crucial tools for business users to gain more insights from customers. However, textual data from customers such as comments are usually short. Thus, there are two main issues in short text categorization: (1) insufficient contextual information and (2) noisy data due to misspellings. Recently, there is a prior attempt to propose a deep learning approach for short text categorization by expanding semantic using word embeddings clustering via an algorithm based on density peaks searching. Since the number of words in word embeddings is usually large, the clustering algorithm does not scale with the size of data set, thus demanding unacceptable computational cost. In this paper, we aim to propose a fast short-text categorization framework. Rather than using a CNN with k-max pooling layer, we propose to use a faster version of Convolutional Neural Network (CNN) called “multichannel CNN.” To speed up the semantic expansion process, we propose to employ mini batch K-Means++, which is considerably faster and scales well with the size of data set. Furthermore, we also introduce an additional preprocessing step to increase vocabulary coverage rate on word embeddings. We conducted experiments on four public data sets: Google Snippets, TREC, MR, and Subj. The results showed that the proposed framework does not only improve an accuracy on all data sets, but also reduces computational costs from several days to a few hours.