Naïve Bayes
Fred Nwanganga, Mike Chapple · 2020
This chapter introduces a new classifier called naive Bayes, which uses a table of probabilities to estimate the likelihood that an instance belongs to a particular class. It presents a spam-filtering example to illustrate how the naive Bayes classifier can be used to label unseen emails based on how similar prior emails were labeled. The chapter discusses the basic principles of probability, joint probability, and conditional probability. It explains how the naive Bayes classification approach works and how that differs from classical Bayesian methods and how to build a naive Bayes classifier in R and how to use it to predict the class values of previously unseen data. The strengths and weaknesses of the naive Bayes method are also examined. The chapter also presents a dataset of more than 1,600 email messages, labeled as either "ham" (legitimate messages) or "spam" (unsolicited commercial email) for the analysis.