Knowledge Extraction from Distributed Heterogeneous Data Sources
Raji Ramachandran, K S Abhijith, J Karthik · 2024
The growing volume and complexity of unstructured and semi-structured data pose a significant challenge in extracting meaningful and relevant information. Information Extraction (IE) emerges as a powerful technique to transform unstructured data into structured format, facilitating further analysis and knowledge discovery. IE techniques can effectively extract information from unstructured data in the form of entities, relationships, objects, and events. Our objective is to develop a method for extracting or filtering out the most relevant information from data sources which are heterogeneous in nature as well as they are distributed over many systems and hence build a unified schema based on the collective knowledge extracted from the data sources. Several existing methods have been developed to extract knowledge from heterogeneous datasets. We focuses on datasets in the medical field that are distributed over multiple systems. Specifically, we aim to extract valuable information from patient records, including patient details, medical conditions and medications consumed. The unified schema can be utilized to study a new patient's condition and assist doctors in making a pre-examination of a patient by querying it.