Hybrid Nature-Inspired Based Oversampling and Feature Selection Approach for Imbalance Data Streams Classification

Monika Arya, Bhupesh Kumar Dewangan, Monika Verma, Mrs. P. Rohini, Anand Motwani, Sumit Kumar Sar · 2023

A big data stream is described using 5 $\mathrm{V}prime$s (Volume, Variety, Velocity, Variability, and Veracity). These characteristics impose various challenges. In many cases, data streams are unbalanced, making traditional data mining approaches impossible to employ. The standard data mining approaches are not suitable for imbalanced data streams for achieving analytical efficiency because they require periodic analyses, but big data requires real-time analytics. Additionally, the induction model must be re-run and rebuilt each time to add up the most recent data. Mining these unique streams, on the other hand, is one of the most intriguing research areas. Deep learning (DL) algorithms were developed to increase classification performance for issues requiring large data sets with varying types and characteristics. Feature selection (F.S.) is a critical stage in any classification application. F.S. entails eliminating superfluous and redundant characteristics, resulting in a prediction model that is more efficient, interpretable, and fast. While complete solutions are available for F.S., managing massive data streams that need instantaneous processing is challenging by its own nature. This work provides a hybrid metaheuristic strategy with FFSMOTE for oversampling imbalanced data stream and Honey Bee algorithm for feature selection. Hybridization aims to improve feature selection processes by combing the advantage of both the algorithms. Furthermore, the Ensemble Deep classifiers are used for data stream classification.

Read the paper · More papers on PaperTik