Spam Text Messages Detection Using Multilingual Embedding
Rohit Kumar Sachan, Vanya Tiwari, Nabh Raghav, Prince Rajpoot · Procedia Computer Science · 2025
As the prevalence of Short Message Service (SMS) spam continues to escalate globally, there is an urgent need for robust and adaptable detection mechanisms. This paper proposes an approach to SMS spam detection utilizing multilingual embeddings. By leveraging advanced techniques like multilingual model embedding, sentence transformer, and AutoML tool, we construct a machine learning model to discriminate between spam and ham messages. By employing multilingual embedding, our model gains the ability to generalize patterns indicative of spam across a wide array of linguistic contexts. Our experimental results demonstrate the improved accuracy of our approach compared to other reported models of state-of-the-art works. The proposed model exhibits 98.85% balanced-accuracy, 99.0% accuracy, 98.0% precision, 93.0% recall, and 95.0% F1 score. This result shows its potential as a valuable approach in the ongoing battle against SMS spam. Furthermore, the model exhibits robustness in detecting spam messages across languages not encountered during training, showcasing its potential for real-world application in multilingual environments.