A Study on Data Profiling Based on the Statistical Analysis for Big Data Quality Diagnosis
Won-Jung Jang, Jong-Yoon Kim, Bum-Taek Lim, Gwang-Yong Gim · International Journal of Advanced Science and Technology · 2018
The volume of digital information produced and distributed globally is expected to be around 90 zetabytes (ZB) by 2020.In the era of the Fourth Industrial Revolution, all things are connected to the Internet and various big data are being produced explosively.Industries are increasingly demanding big data to improve productivity of products, services, and factories using artificial intelligence, but systematic studies on the quality of industrial sites and public data are lacking.Artificial intelligence requires a good overview of data quality before demanding high quality data and taking advantage of big data.The purpose of this study is to propose a data profiling model using statistical analysis techniques to derive attributes for big data quality diagnosis.To do this, the R package and the Delhi Weather data set registered with Kaggle for empirical studies are used in this study.This study calculated property weights and attribute corrections by using empirical methods of statistical analysis for all attributes, and confirmed that the performance comparison study model is superior to the value (accuracy) error rate calculation model for attributes derived from the research model.It is expected that data profiling can be performed in a more scientific way rather than relying on the subjective judgment of the performer in the big data quality diagnosis.