Emotion and Phrase-Based Patterns in Smishing: A Feature-Driven Detection Framework
Natalia Krawczyk, Barbara Probierz, Jan T. Kozák · Procedia Computer Science · 2025
Smishing, or SMS-based phishing, remains a significant threat to mobile users due to its use of concise and emotionally manipulative language. These messages often rely on psychological cues and high-risk phrases that are not fully captured by traditional feature engineering. This study proposes an improved smishing detection framework that adds emotion-based labels and phrase-level risk indicators to embedding-based models. We used GPT-4.5 to annotate emotional dimensions. We also extracted targeted phrasal indicators based on known smishing patterns. Experiments on a unified dataset of over 22000 SMS messages show that combining emotion-based and phrase-level features consistently improves classification accuracy. For example, in Logistic Regression with TF-IDF, accuracy increased from 0.9186 to 0.9221 after adding both types of features. Gradient Boosting with TF-IDF showed the largest performance gain, while Random Forest achieved the highest absolute accuracy in this setting (0.9432). For Word2Vec embeddings, similar patterns were observed: Logistic Regression improved the most (from 0.8721 to 0.8817), while XGBoost reached the highest overall accuracy (0.9447) with sentiment-enhanced input. However, in some cases, risk-based features led to a slight drop in performance, suggesting potential feature noise in highly optimized or shallow models. These findings indicate that combining emotion-driven and pattern-based features with classical embeddings enhances phishing detection performance, particularly in short messages where traditional methods may lack context awareness.