Derivation of Document Vectors from Adaptation of LSTM Language Model
Wei Li, Brian Kan-Wing Mak · 2017
In many natural language processing tasks, a document is commonly modeled as a bag of words using the term frequency-inverse document frequency (TF-IDF) vector.One major shortcoming of the TF-IDF feature vector is that it ignores word orders that carry syntactic and semantic relationships among the words in a document.This paper proposes a novel distributed vector representation of a document called DV-LSTM.It is derived from the result of adapting a long short-term memory recurrent neural network language model by the document.DV-LSTM is expected to capture some high-level sequential information in a document, which other current document representations fail to do.It was evaluated in document genre classification in the Brown Corpus , the BNC Baby Corpus, and the Penn Treebank Dataset.The results show that DV-LSTM significantly outperforms TF-IDF vector and paragraph vector (PV-DM) in most cases, and their combinations may further improve classification performance.