Machine learning-based smishing detection using fuzzy logic and TF-IDF feature engineering

Santosh Kumar Birthriya, Priyanka Ahlawat, Ankit Kumar Jain · Franklin Open · 2026

Mobile communication security is increasingly threatened by smishing messages, necessitating advanced detection techniques to protect users from fraudulent and malicious content. This paper presents a hybrid approach that combines Term Frequency–Inverse Document Frequency (TF-IDF) with fuzzy membership–based linguistic and structural features to enhance smishing messages classification. The feature extraction process includes word count, punctuation usage, message length, sentiment polarity, capitalization patterns, and digit frequency. Fuzzy membership functions encode these attributes as gradual values rather than fixed thresholds, improving adaptability to evolving smishing patterns. These fuzzy features are concatenated with TF-IDF vectors to form a comprehensive representation that captures both semantic and stylistic characteristics. The proposed framework is evaluated on a dataset of 6,119 SMS messages, comprising 5,574 messages from the SMS Spam Collection v.1 and an additional 545 smishing messages from the Smishtank repository. Experimental results demonstrate that the proposed model achieves up to 99.10% accuracy, 99.30% precision, and 94% recall, outperforming existing methods such as SVM (97.40%) and Random Forest (98.15%). Furthermore, the incorporation of fuzzy membership concepts enhances adaptability to diverse smishing patterns, reduces false alarms, and improves the overall robustness of the classification framework.

Read the paper · More papers on PaperTik