BERT-Based Natural Language Processing System for Online Social Information Retrieval on Illness Tracking
Mohit Chowdhary, Poonam Nandal · 2023
With the growing rise of digital platforms such as facebook, twitter and others that attract billions of people around the world, it has gained a lot of traction as a source of unstructured data. The volume of unstructured data is expected to expand at a 60 percent annual rate in the future years. The emergence of social media has facilitated the expansion of unstructured data. Thanks to social media, everyone may now express and share their thoughts and opinions on a wide range of topics. Infodemiology is one such popular and rapidly expanding information retrieval method (i.e., information epidemiology). As per the concept, a set of approaches that study data, particularly health data on the internet, for the purpose of public health studies and policy. Web data is analyzed to detect disease outbreaks faster than traditional surveillance in infoveillance (also known as information surveillance). This research has created a prototype software model for extracting spatiotemporal diseases from diverse data sources in this vein. To identify prospective sickness instances, the proposed method gathered tweets and other postings and used NLP semantic algorithms with transformers called BERT. The prototype system is tuned to improve disease detection using rules given by dissimilar semantic similarity algorithms. The prototype’s output is a visual representation of likely disease forms and spread, which is a straight IE produce from Twitter and other sites.