An integration of corpus applied Linguistics with deep learning for Arabic Spam and Ham Tweets Detection model
Monir Abdullah, Najla I. Al-shathry, Muhammad Swaileh A. Alzaidi, Ali Abdulaziz Aljubailan, Abdulbasit A. Darem, Hanan Al Sultan, Ahmad A. Alzahrani, Abdulwhab Alkharashi · Journal of Radiation Research and Applied Sciences · 2025
Twitter users, comprising organizations and individuals, face the main common difficulty of handling spam tweets. This issue continues to increase owing to numerous social media platform users and rising popularity. Spammers will use this appealing attention to post videos, texts, and images enclosing deceptive links, which might be employed to propagate negative marketing, sarcasm, rumours, and viruses and capture the consumer's information. Detection of Spam is the current cybersecurity and natural language processing (NLP) research field. Numerous studies have been carried out on this esteem, but they mainly emphasize the English language. Conversely, Arabic tweet spam recognition still has a long manner of driving, primarily highlighting the various dialects except for modern standard Arabic (MSA), so, in the tweets, the standard language is rarely utilized. The condition needs an efficient, automatic, robust Arabic spam tweet recognition technique. This study presents an Integration of Corpus Linguistics with a Deep Learning Model for Arabic Spam and Ham Tweets Detection (ICLDL-ASHTD) technique. The main aim of the ICLDL-ASHTD technique is to classify and detect Arabic spam/Ham tweets. In the ICLDL-ASHTD model, the data preprocessing is primarily performed to transform the input tweets into a standardized format. Furthermore, the Term Frequency Inverse Document Frequency (TF-IDF) method is utilized as a feature extractor to produce a feature vector. The ICLDL-ASHTD model utilizes a self-attention bi-directional long short-term memory (SA-BiLSTM) approach for Arabic ham and spam tweet recognition. Finally, the manta ray foraging optimization (MRFO) approach is utilized for hyperparameter tuning. The ICLDL-ASHTD technique is validated through simulation using Arabic tweet datasets. The experimental validation of the ICLDL-ASHTD technique portrayed a superior accuracy value of 98.17 % compared to recent methods.