Role of data normalization in k-means algorithm results
Lamia AbedNoor Muhammed · AIP conference proceedings · 2023
The quality of data is an important issue in data analysis process. While raw data comes from real word may have different problems in its quality; missing data, outliers. Also numeric data may be come in different scales because it is measured using different units according to features that maintains. The less data quality can give down the analysis results. This problem is more obvious in data clustering, however, data clustering is unsupervised task that based on the data. One of probabilistic solutions to overcome the problem of using different scales in measuring data, normalization technique has been used to scale the data in a unified ones. So the produced data have same domain. Different normalization techniques can be applied. This paper shed light on the performance of normalization with one of the most famous clustering algorithms that is k-means clustering algorithm. We chose this algorithm because it is very sensitive to data quality, so the results can be distinguished according to this problem. The practical work used four specific data sets. The work covers the application of data without normalization. Then the different normalization techniques would applied on these data separately. To perform the comparison, accuracy metric would be done to evaluate the performance. The results were obtained revel the important of the normalization in increasing the accuracy of data sets; but in diverse directions according to data used. However, data sets respond very well to specific normalization technique, in opposite side, there is shut down in accuracy with other. So this results must be taken more attention to work in this field in selection the suitable technique to improve the results. Not always give the same tool for applying with different data set. At last, the paper was organized in: introduction, methods and materials, practical work and results, discussion and conclusion.