Exploring Data through the Topological Lens: a Topological Machine Learning Pipeline for Data Analysis
Claudia Caudai, Sara Colantonio, CONTI, FRANCESCO, Mario D’Acunto, Davide Moroni, Maria Antonietta Pascali · HAL (Le Centre pour la Communication Scientifique Directe) · 2024
The advent of deep learning has led to astonishing advancements in data analysis, enabling a shift from handcrafted features to automatically identifying features through representation learning. However, the family of learnable functions is constrained by the specific machine learning paradigm, such as convolutions combined with non-linear activation functions in the popular case of Convolutional Neural Networks (CNNs). Although in principle these paradigms can capture arbitrary patterns in the data,serving as universal approximators, other classes of descriptors could be integrated into machine learning frameworks.Topological invariants are known from mathematics to offer crisp and computable descriptors suitable for differentiating spaces. When applied to real-world data, these descriptors might seem too rigid. However, thanks to the theory of persistent homology, it is possible to use them to conduct intrinsically multiscaleanalysis and propose a Topological Machine Learning (TML) pipeline.The development of a TML pipeline involves two crucial steps that strongly influence the performance of the pipeline: (i) the choice of the filtration that associates a persistence diagram with digital data; and (ii) the choice of the representation method for the persistence diagrams, which often relies on severalparameters. We developed a pipeline that associates persistence diagrams to digital data, via the most appropriate filtration for the type of data considered. Using a grid search approach, this pipeline determines representation methods and parameters, which are optimal for the classification task assigned. We assessed the performance of our pipeline, and in parallel, we compared the different representation methods, on popular benchmark datasets. In real-world problems, the pipeline has been exploited in the medical domain for the classification of Raman spectroscopy: the TML pipeline achieved convincing results for the grading of chondrosarcoma while in a second work it has been used to distinguish the Alzheimer disease from other neurodegenerative pathologies. This line of research has two main goals: first, to offer an easy-to-use pipeline for data classification that integrates persistent homology with machine learning; second, to explore the theoretical reasons why certain combinations of filtrations and topological representations work better for specific datasets and tasks.