Study on Missing Values and Outlier Detection in Concurrence with Data Quality Enhancement for Efficient Data Processing

Feby A. Vinisha, L. Sujihelen · 2022 4th International Conference on Smart Systems and Inventive Technology (ICSSIT) · 2022

Data analytics is the process of analyzing raw data to make predictions and derive conclusions. This process involves collecting and organizing data to discover hidden patterns and draw insight into the data. The methods and approaches of data analytics are automated using various algorithms and mathematical formulas. Data analytics provides real-time and actionable perceptions on data that enable more accurate and prompter decision-making. During the data acquisition phase, missing values and outliers are encountered that affect the model's reliability. Missing values are the data missing in a dataset, more common on large datasets that arise due to information loss, dropout, or non-response of participants. Missing values affects the accuracy of the result and also may lead to biased results. Outliers are the abnormal values that happen to have deviated from the normal distribution pattern of data distribution. Outliers are extreme, unrealistic, extremely big, or small values in a dataset that arise due to manual errors like participant response errors and data entry errors. Outliers also affect the accuracy of the results and lead to over or underestimated resultant values. As missing values and outliers degrade the performance of the analytical data models, various research works have focused on finding and handling such values. This paper reviews the various aspects of missing values and outliers in the preprocessing phase of data analytics to enhance the accuracy of the data model.

Read the paper · More papers on PaperTik