Tensor factorization for cross lingual entity co-reference resolution in the linked open data
Melkamu Beyene, Pierre-Édouard Portier, Solomon Atnafu, Sylvie Calabretto · 2015
The main objective of this research was to identify co-referent entities located in several linked open data (LOD) sources that are described in various natural languages. The problem is approached from two perspectives. First, we do a multi-scale analysis of the RDF graph to discover structural similarities of entities. This was implemented as a tensor decomposition of the RDF graph with each predicate corresponding to a horizontal slide of the tensor. Hereafter, we used the term "structural evidence" to refer to the result of this analysis. Second, for each entity, we associated textual data coming from the Web of documents. Thus, after some preprocessing (viz. removing empty words, applying a weighting scheme such as tf-idf, ...), we represented each entity in a high dimensional space with each dimension corresponding to a term. Next, through a Singular Value Decomposition (SVD), we find a subspace such that the sum of squared distances from the original space to the sub space is minimized. This dimensionality reduction allows us to find language independent similarities between entities. Hereafter, we use the term "textual evidence" to refer to the result of this analysis.