Large–Scale Knowledge Graph Embeddings in Apache Spark
Bedirhan Gergin, Charalampos Chelmis · 2024
RDF2Vec has emerged as a popular method for unsupervised feature extraction from RDF graphs. However, RDF2Vec cannot handle large graphs efficiently, due to (i) the size of RDF graphs or intermediate results being prohibitively large, and (ii) the high computational complexity associated with graph walks of increasing breadth and depth, which makes their processing difficult, if not impossible, on a single machine. We address this limitation by introducing SERE, a scalable and distributed framework for unsupervised embeddings computation on large–scale Knowledge Graphs. SERE is open–source, well–documented, and fully integrated into the SparkKG–ML Python library. Our experiments demonstrate that SEREis able to compute embeddings over Knowledge Graphs with millions of edges within hours, and is up to 7 times faster than RDF2Vec even for small Knowledge Graphs, all while achieving comparable accuracy to RDF2Vec in a benchmark classification task.