Secure NoSQL Based Medical Data Processing and Retrieval

Roshan Ramprasad Shetty, Akalanka Bandara Mailewa, Susan A. Mengel, Lisaann S. Gittner, Ravi Vadapalli, Hafiz Khan · 2017

The transdisciplinary big data medical research of the Exposome project discussed in this paper can be best described by the old adage 'Finding a needle in a haystack', in this case, of Excel .csv or other types of disparate files from various service providers, such as the US Census Bureau and the Center for Disease Control and Prevention of the US Department of Health and Human Services. The Exposome project aims to bring together such data files from different sources to draw previously unknown insights into medical and other types of issues, such as cardiovascular disease and infant mortality. Data from these and other providers, however, continues to grow with new data added frequently; so, the process of finding the correct or desired data is a quite cumbersome and time-consuming task. The data may also be unstructured causing relational databases to be less optimal for handling and processing such enormous data volumes. Thus, NoSQL databases offer a promising means to organize and allow parallel access to the data as presented in this paper. The aim is to provide a retrieval system that deals with disparate files and unstructured data while providing an easy and efficient solution for data selection to the researchers. The paper discusses various processes that can be used and optimized to eliminate critical challenges haunting big data processing. The key features of the system include: auto data dictionary creation, conflict resolution, null data handling, and an assistive user interface.

Read the paper · More papers on PaperTik