AutoML – Optimal K Procedure

Oded Koren, Michal Koren, Amit Sabban · 2022

Clustering procedures are an important method in the AI domain. They help improve machine learning capabilities and allow the expansion and enrichment of a variety of services, functionalities and information for organizations, decision-makers, recommendations, predictions, etc. One of the major challenges of the k-means algorithm is selecting an optimal K cluster for the given classified datasets. Our research defines, implements and examines an AutoML procedure that combines numeric and categorical datasets, which are divided into subsets of numeric values, analyzed and compared. Then, via four different scores, the procedure selects an optimal k value according to the scores, after evaluating the square(n) observations and integrating the optimal k into the datasets, with an additional labeling value. The research presents experiments of the end-to-end AutoML procedure, which is implemented using two datasets. Furthermore, it elaborates on the advantages of this innovative procedure that analyzes, selects and implements the optimum k for a dataset.

Read the paper · More papers on PaperTik