Drug Side Effect Frequency Mining over a Large Twitter Dataset Using Apache Spark
Dennis Hsu, Melody Moh, Teng-Sheng Moh, Diane Moh · Apple Academic Press eBooks · 2020
Despite clinical trials by pharmaceutical companies as well as current Food and Drug Administration reporting systems, there are still drug side effects that have not been caught. To find a larger sample of reports, a possible way is to mine online social media. With its current widespread use, social media such as Twitter has given rise to massive amounts of data, which can be used as reports for drug side effects. To process these large datasets, Apache Spark has become popular for fast, distributed batch processing. To extract the frequency of drug side effects from tweets, a pipeline can be used with sentimental analysis-based mining and text processing. Machine learning in the form of a new ensemble classifier using a combination of sentiment analysis features to increase the accuracy of identifying drug-caused side effects. In addition, the frequency count for the side effects is also provided. Furthermore, we have also implemented the same pipeline in Apache Spark to improve the speed of processing of tweets by 2.5 times, as well as to support the process of large tweet datasets. As the frequency count of drug side effects opens a wide door for further analysis, we present a preliminary 234 study on this issue, including the side effects of simultaneously using two drugs, and the potential danger of using less common combination of drugs. We believe the pipeline design and the results present in this work would have helpful implication on studying drug side effects and on big data analysis in general. With the help of domain experts, it may be used to further analyze drug side effects, medication errors, and drug interactions.