MapReduce based parallel data processing for drug-drug interaction prediction

Faran Wei · 2016

Prediction of drug-drug interactions (DDIs) can prevent unexpected adverse drug events (ADEs).ADEs can damage people's health and bring about economic losses to the society.Existing methods usually adopt the dataset of the adverse drug reports (ADRs) to build the DDI prediction statistical model; however, the ADRs dataset contains billions of records and the U. S. government releases new ADRs quarterly so that the data volume keeps growing.It is time consuming to clean these data and extract useful features from ADRs, which delay the development of effective DDIs prediction models.In this paper, a parallel processing framework based on MapReduce model is proposed.The MapReduce model is utilized to extract names and adverse reactions of drugs and count their frequencies based on ADRs, which can improve the efficiency of data cleaning and feature extraction of the Food and Drug Administration (FDA)'s adverse drug reports.The parallelization of the statistical screening and processing method of ADRs are implemented on the Hadoop cluster.Experimental results show the processing of FDA's adverse drug reports can be achieved accurately by this method and the speed of processing can be improved effectively by using the proposed parallel computing framework.

Read the paper · More papers on PaperTik