Synthetic Data for Scam Detection: Leveraging LLMs to Train Deep Learning Models
Pitipat Gumphusiri, Tuul Triyason · 2024
This paper presents a novel approach to training scam detection models using synthetic data generated by Large Language Models (LLMs). We propose single-agent and multi-agent methods for data generation and train six deep learning architectures-LSTM, BiLSTM, GRU, BiGRU, CNN, and BERT-to classify conversations as scam or non-scam. Our experiments demonstrate that models trained on synthetic data achieve high accuracy on both generated test sets and real-world scam conversations. The models perform well even with limited conversation turns and when analyzing only the suspect's messages, indicating potential for early scam detection and privacy-preserving applications. Our findings highlight the efficacy of synthetic data in overcoming real-world dataset limitations for scam detection. We make the dataset and trained models publicly available to facilitate further research and development in this critical area of fraud prevention.