How to Tackle Fake Medical Content with Clustering and Transformers

Radu Răzvan Slăvescu, Marina Bianca Trif, Anca Nicoleta Marginean, Kinga Cristina Slăvescu · 2024

Spreading fake medical information / news online might threat the public health, e.g., by decreasing the vaccination rates or the level of trust in the healthcare process. The system presented in our paper aims to address this by determining in an automatic manner whether a text either contains or implies some claims known to be false. To this end, it relies on the Monant dataset, which comprises medical posts and claims whose veracity has been established by experts in the field. The dataset is used in two ways. First, the set of false claims aimed to be detected was obtained from it. These false claims are clustered, allowing the next steps to be performed against one cluster only, thus reducing the response time. Second, to detect a post's falsehood, this gets summarized using transformers, then the summary is compared similarity-wise with each element of the chosen cluster of false claims, using all-MiniLM-L6-v2 and a variation of BioBERT. A high similarity score is interpreted as the presence of the false claim in the summary. The results obtained from the experiments demonstrate, through manually checked examples, that the system manages to identify about 77% of the false claims within the text at hand (with an F1-value of 0.73).

Read the paper · More papers on PaperTik