Application of K-means Clustering Model to XRD Experimental Data in the Korea Plateau

Ju Young Park, Sun Young Park, Jiyoung Choi, Sungil Kim, Yuri Kim, Bo‐Yeon Yi, Kyungbook Lee · Economic and Environmental Geology · 2024

Mineral composition used to identify the sedimentary environment can be obtained through X-ray diffraction (XRD) analysis.However, due to time constraints for analyzing a large number of samples, a machine learning-based mineral composition analysis model was developed.This model demonstrated reasonable reliability for samples with usual compositions but showed poor performance for unusual samples.Consequently, a clustering model has recently been developed to classify the unusual samples, allowing experts to handle.The purpose of this study is to examine the applicability of the clustering model, developed using XRD data from the Ulleung Basin in previous study, using samples from different regions.Research data consist of intensity profile from XRD experiment and its mineral composition analysis for a total of 54 sediment samples from the Korea Plateau, located northwest of the Ulleung Basin.Because the intensity of samples in the Korea Plateau comprises 7,420 values (3.005-64.996°),differing from 3,100 values (3.01-64.99°) of samples in the Ulleung Basin, linear interpolation was used to align the input feature.Then, min-max scaler was applied to intensity profile for each sample to preserve the trend and peak ratio of the intensity.Applying the clustering model to the 54 preprocessed intensity profiles, 35 samples and 19 samples were classified into expert and machine learning groups, respectively.For machine learning group, false positive was zero among the 19 samples.This means that the clustering model can increase reliability in when mineral composition from machine learning model because unusual sample did not belong to the machine learning group.For the 35 samples in expert group, the 31 samples were classified as false negative (FN).It means that although machine learning model can properly analyze these samples, they were assigned to expert group.However, when these FN samples were analyzed using machine learning based composition analysis model, a high mean absolute error of 2.94% was observed.Therefore, it is reasonable that the samples were assigned to expert group.

Read the paper · More papers on PaperTik