Performance Analysis and Classification of Class Imbalanced Dataset Using Complement Naive Bayes Approach

Bhaskar Marapelli, Sreedevi Kadiyala, Chandra Srinivas Potluri · 2023

Data mining technology is essential to all of the major engineering professions in the modern world. Big data is a developing trend. The amount of data is increasing exponentially these days. The complexity of this data makes it difficult to analyze, store, and process. To manage this massive amount of data, we need a system that can extract the intricate representation of the data. For the purpose of overcoming this issue, the complement naive bayes (CNB)algorithm is employed. One of the most popular classification methods is the complement naive bayes algorithm; however, a new data pre-processing method has been proposed that can deal with imbalanced dataset concerns. A balanced dataset can be classified as the number of examples of some class is equal to or lesser than the number of examples belonging to other classes and the dataset is evenly distributed. Imbalanced datasets can be classified as the number of examples of some class is higher than the number of examples belonging to other classes and datasets are not evenly distributed in addition to this, an Inherent issue will occur in an imbalanced dataset. In the suggested article, the dataset may be distributed uniformly and the underlying problems can be effectively resolved by using the CNB technique. There are some libraries available to process huge data such as hive permits, Kafka, Hadoop, and mahout. Hadoop is utilized in this study to determine the class and attribute probabilities. Using training datasets, experiments were conducted, and reviews were categorized as a working model and a failure model. The Performance of the classes' working model and failure model were compared according to working time, accuracy, memory space, quality, speed, and security.

Read the paper · More papers on PaperTik