Composed clustering of non-relational data with mixed types using K-modes and K-prototypes algorithms
Makrem Jannadi, Khaled Ben Driss · 2020
The current available structured data for the majority of AI projects contains mixed categorical and numerical types. This mixture of data types drives the researchers to continually invent and develop new algorithms capable of handling these issues. K-Modes and K-prototypes are good examples of these innovative algorithms for clustering purposes. However, many fields are generating non-relational data that offers flexible schema and enables faster computing. The problem with this data is that it needs an advanced preprocessing step in order to make it ready to use by clustering algorithms. We propose a composed clustering that can fix the issue of document-based categorical attributes, by using K-Modes to cluster the modalities of these attributes and standardize them, then applying K-Prototypes for clustering individuals of the dataset using all numerical and preprocessed categorical attributes. We tested our approach on a real-world database used by human resource agents to cluster applicants based on information extracted from their resumes, with results showing good performances of our approach.