Towards Extracting Drug-Effect Relation from Twitter: A Supervised Learning Approach
Fan Yu, Melody Moh, Teng-Sheng Moh · 2016
Advancements in social media technology have resulted in the booming of massive public data. The availability of these huge data sets offers numerous research opportunities for deriving meaningful cause-effect relationships for many applications. One important application domain is the cause of side effects of drugs. In this paper, we applied supervised learning to extract useful cause-and-effect information related to drugs from Twitter. To filter out unrelated information and to increase the accuracy of classification, a spam filter and a preprocessing procedure have been developed. Validation experiments were performed using a manually labeled data set based on streamed tweets collected continuously on Twitter in real-time for 48 hours, and exploiting six different supervised machine-learning classifiers. Results have shown that these classifiers have achieved up to 77% accuracy in identifying drugs' cause-effect relations on Twitter data. This result has shown a positive feasibility for collecting drug side effect information from Twitter. The proposed method may be applied to other areas such as food, beverages, and other daily consumer products for finding their side effects and people's opinions concerning them.