Facilitating Knowledge Graph Analysis – Acquisition and Large-Scale Analysis of Topological Graph Measures
Matthäus Zloch · Univ. Duesseldorf: Duesseldorfer Dokumenten- und Publikationsserver · 2021
In today’s Web, the most common model for structuring knowledge and making it machine-readable is the knowledge graph. In this model, vertices represent Web entities that encode real-world objects as URIs (Uniform Resource Identifiers); edges are labeled, and represent relationships between these entities, which are modeled by knowledge-domain-specific vocabularies and predefined schemas. The topology of knowledge graphs differs fundamentally from other topologies, for example those of computer networks or social graphs. This is because, first, knowledge graphs contain hierarchical (typed) as well as transversal relationships between vertices. Second, the shape of the graph topology is significantly influenced, on the one hand, by knowledge-domain-specific vocabulary usage defined by particular schemas and, on the other hand, by the inconsistent modeling habits of researchers and modeling tools. Analyzing and understanding the distinct topology, and employing meaningful measures for the appropriate characterization of knowledge graphs is crucial, and can guide and inform the development of, for example, profiling tools, benchmarking solutions, efficient data structures and indexes, and compression techniques. Traditional measures known from network science inadequately capture the semantics that knowledge graph topologies entail. Therefore, it is of central importance to provide appropriate tools for the analysis, and proper measures for the characterization of knowledge graphs. The present cumulative dissertation is motivated by this. It makes three scientific contributions, each of which constitutes one part of the thesis. The first part of the thesis introduces and describes a software framework that consolidates third-party tools for the acquisition and preparation of knowledge graphs in order to enable graph-related tasks on their topology. We perform a large-scale analysis of 280 knowledge graphs from nine knowledge domains provided by the Linked Open Data (LOD) Cloud, and we calculate 54 different graph measures with this tool. The analysis results and the processed graph objects are available to the research community for further processing. Building on this, the second part of the thesis deals with the investigation of commonly used measures from network analysis as well as measures that have been specially introduced for the characterization of RDF knowledge graphs. We examine them in terms of their relevance and meaningfulness for generating concise descriptions of knowledge graph topologies. In particular, we seek to find measures that have the capacity to discriminate graphs from other knowledge domains in order to reveal knowledge domain specificities and derive corresponding implications for existing solutions. To this end, we employ various statistical methods and a state-of-the-art machine learning classification model. In the third and final part of this thesis, we employ our framework introduced earlier to propose solutions in other research areas of knowledge graphs. We deal with database benchmarks for knowledge graphs and address the criticism that RDF benchmarks deliver less reliable results due to the usage of synthetic queries for runtime measurements. To this end, we propose a functionality of our framework to leverage programmatic graph representations from knowledge graphs to generate application-specific queries based on real-world data. Furthermore, we present a flexible business use case-driven approach, which allows to assess response times of database queries more reliably by means of building query groups. This thesis is based on published papers submitted to high-ranked international peer-reviewed open access journals, international conferences, and workshops in the research area of Semantic Web technologies. As a commitment to open science, all code and resources have been published as open source projects under MIT license on popular code and data hosting platforms.