A Scalable Approach to Aligning Natural Language and Knowledge Graph Representations: Batched Information Guided Optimal Transport
Alexander Kalinowski, Deepayan Datta, Yuan An · 2023
The triples of a knowledge graph (KG), comprised of a subject-object pair and a predicate linking the two, share common structural properties to natural language (NL) sentences. Recent advances in data compression via neural networks, specifically low-dimensional embeddings, have been applied to both data domains (KG and NL), yet little work has been done to explore how these two different representations may be linked together in an attempt to align the knowledge of one domain to another. In this work, we develop an unsupervised deep learning methodology for such alignment using tools of optimal transport and the Wasserstein distances between the respective representations. Due to the size and complexities of each embedding space, we introduce a novel process to improve scalability and reduce training time we call Batched Information Guided Optimal Transport (BIG-OT). We experiment on two datasets (NYT-Freebase and Wikidata) without using the supervision of aligned pairs and show that our solution can outperform the current benchmarks while only operating on sub-samples of each respective space, thus providing gains in algorithmic efficiency. Our algorithm scales to Big Data use cases regardless of the baseline embedding methodologies utilized, and we additionally show a simple procedure to improve those embeddings again leads to performance gains.