Predictive analytics on Electronic Health Records (EHRs) using Hadoop and Hive
Haritha Chennamsetty, Suresh Chalasani, Derek Riley · 2015
Healthcare industry is providing massive amounts of patient data. The need for parallel processing is apparent for mining these data sets to provide personalized medicine or advice to patients. An EHR data management system is essential to provide insights and predict outcomes from past patient data. In this paper, we present an EHR data management system to process massive amounts of healthcare data. The system is built on Hive, which is scalable and dynamic compared to traditional data warehouses. Patient data can be uploaded to Hive from a variety of sources like flat files, web pages, real-time applications and databases. Unlike traditional data warehouses, used for transaction processing and analytics, Hive is used for analytics only. The data can be easily sent to Reports application to generate graphs and charts from the Hive data warehouse. The graphical charts are useful for doctors and researchers to understand and propose medications based on evidence from a large number of past patient records. The predictive analysis is helpful to treat patients using specific medications, based on a number of factors such as lifestyle, family history, smoking habits, and health conditions such as blood pressure and diabetes.