Network representation learning with rich text information
Cheng Hong Yang, Zhiyuan Liu, Deli Zhao, Maosong Sun, Edward Yi Chang · 2015
Representation learning has shown its effectiveness in many tasks such as image classification and tex-t mining. Network representation learning aims at learning distributed vector representation for each vertex in a network, which is also increasingly rec-ognized as an important aspect for network anal-ysis. Most network representation learning meth-ods investigate network structures for learning. In reality, network vertices contain rich information (such as text), which cannot be well applied with algorithmic frameworks of typical representation learning methods. By proving that DeepWalk, a state-of-the-art network representation method, is actually equivalent to matrix factorization (MF), we propose text-associated DeepWalk (TADW). TAD-W incorporates text features of vertices into net-work representation learning under the framework of matrix factorization. We evaluate our method and various baseline methods by applying them to the task of multi-class classification of vertices. The experimental results show that, our method outperforms other baselines on all three dataset-s, especially when networks are noisy and train-ing ratio is small. The source code of this paper can be obtained from