Genetic algorithm cleaning in sequential data mining
Kok Cheng Tan, Daniel Zantedeschi, Amruth N. Kumar, Alessio Gaspar · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 2022
Sequence mining has been receiving increasing attention from the data science community as it helps reveal the underlying patterns in a given phenomenon. For instance, clustering sequences of "actions" taken by users provide understandable trajectories of the events. However, such sequences often include inessential actions recorded during data collection. It is generally challenging to distinguish between noise and information during clustering. This paper proposes using a Genetic Algorithm (GA) to denoise the sequences properly before they undergo clustering. The intent is to improve the clustering quality, as measured by Silhouette Score. We present preliminary results obtained by applying this approach to the data collected by a well-established Intelligent Tutoring System (Epplets.org) used by novice programmers. Our findings reveal that the GA can discover sequence cleanings that improve the quality of the clustering process. These results motivate further research into applying Evolutionary Techniques to data cleaning in sequence mining.