Application and Comparison of Majority Weighted Minority Oversampling Techniques and Random OverSampling Examples Data Balancing Methods on the Vertebral Column Dataset
M Pushpalatha, N. Indira · Journal of Emerging Technologies and Innovative Research · 2021
In recent years, Data science has emerged as most conspicuous multidisciplinary field. The raw data that are gathered from multiple sources usually have issues like missing data, imbalance data, scaling, normalization and so on. Hence, practically almost all raw dataset need to be converted to more application suitable form. From these facts, it is apparent that data preprocessing is extremely necessary phase of any data analysis task. Despite the fact that not all preprocessing method need to be applied on a solitary dataset, but one or more could be required to form usable formatted dataset. Like any dataset, the Vertebral column dataset also isn’t always similarly disbursed to the class label. So right here to solve the problem, in this work, two data balancing algorithms namely Majority Weighted Minority Oversampling Techniques (MWMOTE) and Random Over Sampling Examples (ROSE) applied. Furthermore, simulation results are generated using R Software and assessed estimating processing time. As a result, ROSE algorithm takes less processing time than MWMOTE to balance the data.