Metapath-Enhanced Language Model Pretraining on Text-Attributed Heterogeneous Graphs
Shangheng Chen, Quan Fang, Shengsheng Qian, Changsheng Xu · ACM Transactions on Multimedia Computing Communications and Applications · 2025
Text-Attributed Heterogeneous Graphs (TAHGs), which combine text data with various graph relationship information linked to rich semantic entities, are ubiquitous in real-world scenarios. To extract information from TAHGs, a commonly used method is employing Pretrained Language Models (PLMs). However, existing methods are primarily designed for processing text and face challenges when dealing with graph information, leading to two main issues: incomplete context due to graph sampling and weak integration of text and graph information. In this article, we present a new approach named Metapath-Enhanced Language Model Pretraining (MLMP) on Text-Attributed Heterogeneous Graphs. The proposed model starts by gathering metapath information through pre-computed neighbor aggregation using a simple mean aggregator. Subsequently, this gathered metapath information, combined with textual data, is input into a GNN-nested PLM. Here, GNN components at each layer are nested alongside the transformer blocks of PLMs during the training process. We have also developed corresponding pretraining strategies for joint pretraining. The experimental results indicate that our model efficiently captures information within TAHGs. Across benchmark datasets, it consistently outperforms current state-of-the-art methods, demonstrating remarkable effectiveness in tasks such as link prediction and node classification. Our code is available at https://github.com/chensh911/MLMP .