Natural Language Processing-Based Querying Heterogeneous Data Sources Using Integrated Ontology

S. Swathi · Preprints.org · 2023

It is important to deal with data scattered among heterogeneous information sources, which can be structured, semi-structured, or unstructured, to give consumers a cohesive perspective of the data. Information gathering is challenging as a result, and one of the major causes of this is that data sources are developed to support certain applications. A method to simplify this procedure is to use ontology representation as an intermediate step. An ontology represents a knowledge structure that reasonably reflects the complexity of the real world and is ideally used in many industries today. Created local ontologies from heterogeneous data sources (MongoDB and Neo4j) in project phase I using the Cardiovascular (CVD) dataset. Then, the next step is to derive global ontologies using attribute semantic integration through machine learning models and approaches. Finally, implemented twenty more similarity measures for the ontologies and derived the global ontologies. In phase II of this project, the main objective is to implement and integrate a query system with the global ontology system. We implemented this using NLP-based approaches by building SPARQL queries for our ontologies. The system will translate natural language into a well-structured query format.. Built a query component and created SPARQL queries from these natural language questions. Selected the best possible candidate query graph from which we generate the SPARQL queries. Here, the next step is to explore the combination of distinct sentence encoders to provide better latent sentence representations in the query construction. The SPARQL queries generated are fired on the global ontology output, which is then translated into source-specific queries. Therefore, this system allows unified access to heterogeneous data sources through a natural language querying interface.

Read the paper · More papers on PaperTik