Accelerated supervised learning to detect spam using feature selection and Apache Spark architecture

Mehdi Salkhordeh Haghighi, Mahsa Sahebi · 2020

Spam can have fake links and trick users into redirecting to fake pages and stealing users' information on these pages. Machine learning methods are typically used to detect spam to distinguish the spam pattern from normal text. A good way to detect spam is to use machine learning techniques that can detect the pattern of spam. While large volumes of text or email need to be considered, these methods take a long time to make using these methods impractical. A good way to accurately and quickly detect large volumes of spam is to use big data processing platforms such as Apache Spark. In this research, an attempt has been made to use feature selection by meta-heuristic methods to increase the speed and efficiency in spam detection. In this method, the data is entered into Apache Spark for final processing to be detected by spam with high-speed distributed learning. Experiments show that using the feature selection mechanism can increase learning accuracy by 98.66%. The processing speed in Apache Spark is about 5.453 times higher than non-Spark systems.

Read the paper · More papers on PaperTik