Data Preprocess
Md. Sharif Hossen · 2020
A large amount of data is collected every day from different sources. Most of them are unprocessed, which are difficult to analyze and sometimes become useless as the datasets tend to be inconsistent, missing, and noisy. Before using those datasets, the quality must be maintained. Data preprocessing is a step of the knowledge discovery process that ensures the consistency and quality of the data. Data preparation is a compulsory step in data preprocessing, which prepares the useless data in a usable format to analyze in the next step of data mining. There are several techniques in data preparation, e.g., data cleaning, integration, reduction, transformation, normalization, de-noising, and dimensionality reduction, and so on. In this chapter, we will discuss how to measure the quality of data, address missing data, clean the noisy data, and perform transformation on certain variables.