Spam E-mail Detection by Random Forests Algorithm

Bhagyashri U. Gaikwad, Pratap Pandurang Halkarnikar, M. Tech Student · 2013

Spam officially called unsolicited bulk email or unsolicited commercial email, which is rapidly becoming a major problem on the Internet. It is an attempt to deliver a message, over the Internet, to someone who would not otherwise to receive it. These junk emails may contain various types of message such as commercial advertising, quick rich scheme, pornography, doubtful product, illegal service or viruses. Spam increase the load on the servers and the bandwidth of the ISPs and the added cost to handle this load must be compensated by the customers. In addition, the time spent by people in reading and deleting the spam emails is a waste. As a result of this growing problem, automated methods for filtering such junk from legitimate E-mail are becoming necessary. Automatic email spam classification contains more challenges because of unstructured information, more number of features and large number of documents. As the usage increases all of these features may adversely affect performance in terms of quality and speed. Many recent algorithms use only relevant features for classification. This paper described classification of emails by Random Forests (RF) Algorithm. RF is ensemble learning technique. The Random forest is a meta-learner which consists of many individual trees. Each tree votes on an overall classification for the given set of data and the random forest algorithm chooses the individual classification with the most votes. If identified category is 0 then e-mail is marked as non-spam e-mail otherwise if identified category is 1 then e-mail is marked as spam e-mail

Read the paper · More papers on PaperTik