A study on document classification using multiple distributed representations
Koji Takuwa, Tomohiro Yoshikawa, Felix Jimenez, Takeshi Furuhashi · 2017
Document classification is an essential task in digital society. In document classification, it is important how to represent a document. Topical document classification methods represent a document as Bag-of-Words (BOW). It uses only the number of occurrences of each word, so it ignores the semantic meaning of words. Recent years, document classification methods using Word2Vec are proposed and got much attention. Word2Vec is a tool for learning semantic-syntactic relationship among words as word vectors. The word vectors are called distributed representation. The document classification methods using Word2Vec represent a document as the centroid of word vectors in a document. It uses only semantic meaning of each word, so it ignores the number of occurrences of words. In this paper, we propose a new document classification method combining BOW and multiple distributed representations. Different corpus has different words and phrases, so each distributed representation learned from each corpus is expected to have different semantic meaning.