Efficiency Enhancement of Machine Learning Approaches through the Impact of Preprocessing Techniques
Vineeta Gulati, Neeraj Raheja · 2021
A prediction system is affected by numerous factors. The primary factor is the depiction and traits of instance data. In the presence of unnecessary and redundant data or boisterous and deceptive information, it becomes arduous to discover knowledge during the training phase. One of the persistent challenges in data analytics is identifying and rectifying dirty data. The outcomes will be unreliable decisions and inaccurate analytics if an individual fails to do so. It is known by a large mass of people that the machine learning model performances are affected by the quality of data. Consequently, a significant amount of time is spent by the scientists on cleaning data before model training. Data pre-processing is considered the fundamental stage of ML methods by countless researchers. Nevertheless, only a fraction of works have paid emphasis on the consequences of data processing techniques. This paper addresses the serious impact that issues of data pre-processing can have on the generality performance of an ML algorithm and here all the pre-processing techniques like removal of missing values with all the possible methods, data binning, and data normalization, etc. are discussed. After performing all the pre-processing techniques, divide the dataset into training and testing datasets and then apply machine learning algorithms. It is concluded that to achieve a reliable result or better accuracy data preprocessing plays a very important role and to build a machine learning model we must have deep knowledge of all the pre-processing techniques, also where and how to apply them.