Multiclass Imbalanced Big Data Classification Utilizing Spark Cluster
Tinku Singh, Riya Khanna, Satakshi, Manish Kumar · 2021
Because of the massive increase in data collection and storage that has occurred in recent years, big data applications are increasingly becoming the focus of attention. The difficulty of classification with imbalanced datasets is one of the complexities that make extracting meaningful information difficult, and the key impact arises from its existence in a variety of real-world applications. Because of the variety as well as the veracity of such obtained data, big data is impacted by an imbalance of classes. Furthermore, in real-world data applications, samples from one class, which is the core concern, are frequently vastly dominated by samples from other classes. In this study, we have proposed an approach using block-level undersampling and synthetic data point generation to deal with imbalanced big data. Furthermore, the performance of Random Forest and Decision Tree algorithms in dealing with imbalanced datasets in the big data context has been evaluated. Extensive experiments have been performed utilizing Apache Spark Cluster in the development of the different discussion methods. The proposed technique can handle massive datasets while still offering the assistance required to accurately categories classes with a comparatively less number of instances.