Evaluation of clustering approach with euclidean and Manhattan distance for outlier detection
Mukhtar Eri Suhaeri, Alimudin, Anam Javaid, Mohd Tahir Ismail, Majid Khan Majahar Ali · AIP conference proceedings · 2021
This paper presents the results of the seaweed drying data. We compared the two main approaches to clustering analysis between K-Means and (Partitioning Around Medoid) PAM. In agriculture, Clustering application is the method of finding the new insights in different kind of problems. For enhancement of agriculture, machine learning always seems to be helpful in conversion of raw data into useful information. But in case of outlier presence, machine learning algorithms and techniques could not work well. In this paper, the dataset consisted of 1924 observations are taken for the analysis purpose. The dataset included one dependent variable with 29 different predictors. The variable includes the hourly solar radiation, temperature, humidity, and moisture content. The outliers were found in the dataset. The comparison between PAM and K-Means using both Manhattan and Euclidean distance for internal validity results was evaluated on the performance and evaluation of algorithms in case of outliers. The purpose was to evaluate which used the internal clustering validation on the experimental drying process of seaweed data. The results are beneficial to minimize the outliers data and forecasting.