Enhancing Business-Specific Phishing Chat Detection via Few-Shot Learning LLM Augmentation
Saran Hansakul, Nawaporn Wisitpongphan · IEEE Access · 2025
Phishing through one-on-one chat messages with customers poses significant threats to businesses, yet current detections are hindered by a lack of realistic, context-specific datasets. This research proposes using Large Language Models (LLMs) to augment a limited dataset of 100 real-world chat messages from business-specific retail customer service interactions, generating ten times more realistic synthetic messages. GPT-generated realism scoring was then used to filter highly realistic messages. Human surveys confirmed a strong alignment between GPT-generated realism assessments and human evaluators. Machine learning models (Decision Tree, Decision Forest, and Logistic Regression) showed significant performance improvements when trained on the augmented data. Accuracy improved from ~77-90% (achieved using the original limited dataset) to ~97-99%. This confirms that this new proposed augmentation and realism validation using LLMs significantly improves phishing detection, enhancing protection for business-specific retail customer service chat interactions.