Using Classifiers Based on Large Language Models and Naïve Bayes for Domain Specific Text

Артем Ховрат, Volodymyr Kobziev, Vitalii Volokhovskyi, Oleksii Nazarov · 2024

The development of technologies for auto-generation of content in specific domains leads to the aggravation of possible risks associated with falsified information. Currently, the problem of determining the most accurate and fastest algorithm for countering this form of fraud remains open. This article focuses on checking the effectiveness of approaches based on Naive Bayes Classifier and decoder-only Large Language Models to detect the fact of contextual information falsification. Exclude Bayes algorithm, as target models were chosen GPT -40, Gemini Pro and Mistral N eMo. The results of the research conducted on a self-created data set related to news dedicated to Russia's invasion of Ukraine, and a comparison with a previously validated approach based on neural networks, assert the high efficiency of the proposed solutions and the possibility of its implementation as part of anti-fake module for socially oriented systems.

Read the paper · More papers on PaperTik