DocNet: A document embedding approach based on neural networks
Zhonglin Mo, Jianhong Ma · 2018
Embedding texts into vector spaces is a common and fundamental preprocessing. Despite there are several approaches to put documents into vectors, reducing the dimension and improving ability of expression can still be a problem when facing large scale data and sophisticated demand. Distributed dense vector have been shown to be powerful in capturing token level semantics. In this paper, we propose a new method to embed entire documents into vector space using a deep neural network which described as DocNet in this paper. With DocNet, we trained end-to-end learning the vector space and by that we take all the information including semantics into account. Once this space has been produced, tasks such as classification and clustering can be simply done using standard techniques. Our method introduces triplet loss to train. The benefit is vector space can be learned directly so we can control the final dimension of embedding vectors. To demonstrate performance of our method, we built a clustering system compared with several baseline methods. Experiments prove that our approach achieves state-of-art document clustering performance. Furthermore, it proves that complicated clustering or classification demands can be satisfied by our method.