General SimRank: An interpretable similarity measure in heterogeneous information networks

Chuanyan Zhang, Xiaoguang Hong, Yongqing Zheng · Information Sciences · 2026

Measuring object similarity in information networks is a fundamental problem with broad applications, such as recommendation systems and information retrieval. With the increasing prevalence of large-scale heterogeneous information networks (HINs) that consist of multiple types of relationships, developing effective similarity measures has become crucial. However, existing approaches have inherent limitations: (1) Methods designed for homogeneous networks fail to capture edge semantics; (2) Meta-path-based methods rely on expert-defined paths; (3) GNN-based methods lack interpretability. These limitations highlight the need for a global, semantics-aware solution. In this paper, we introduce General SimRank (GSR), an extension of the well-known SimRank, designed for computing interpretable similarity in HINs. GSR refines similarity computation by restricting it to same-typed nodes and ensuring semantic consistency through constrained random walks. We further propose a domain-independent semantic analysis method based on entropy theory to quantify the contribution of different semantic edges to node similarity computation. Additionally, we establish a series of desirable properties of GSR and propose its equivalent formulation as the Constrained Expected Meeting Distance (CEMD) on the graph. Extensive experiments on two real-world HIN datasets validate the effectiveness of GSR, demonstrating its superiority over state-of-the-art methods in both accuracy and interpretability.

Read the paper · More papers on PaperTik