Classification of galaxies using machine learning
Maurizio D’Addona · Zenodo (CERN European Organization for Nuclear Research) · 2021
In this work, I investigate the possibility of finding a data-driven solution to the problem of automatic classification of galaxies using machine learning methods. In the modern scientific contest, the ability to reliably classify distant galaxies is important not only because it allows us to better understand their formation and evolution, but also because it possibly enables us to gain a better insight into the structure formation of our Universe. Due to the lack of a reliable knowledge base and also to minimize human biases, I tested two unsupervised methods on spectra obtained from the Data Release 3 of the Galaxy And Mass Assembly survey (GAMA): namely an Unsupervised Random Forest (URF) and a combination of T-distributed Stochastic Neighbor Embedding (T-SNE) and Density-Based Spatial Clustering of Applications with Noise (DBSCAN). Both algorithms have often been used and validated in various astrophysical contexts but while URF has been successfully used on both star and galaxy spectra (for example Reis et al. 2018 and Baron et al. 2017), T-SNE, to my knowledge, has been only used on spectra of stars (Traven et al. 2017). A sample of approximately 72,000 good quality spectra has been selected from the original set of 166,000 ones available from the GAMA DR3. Each spectrum has been normalized by fitting and subtracting its continuum, then it has been de-redshifted to the rest frame and, finally, it has been rebinned and trimmed to get a set of fluxes in 2250 different wavelength bins ranging from 3700A to 7000A. This range has been chosen to minimize the number of noisy pixels and missing data. The whole sample of spectra has been imputed to fill any missing data. The reduced sample of spectra has been analyzed with URF first, from which however no conclusive result can be extrapolated. In a second experiment, I tried to perform a clustering process with DBSCAN on the sample. However, due to its high dimensionality, no clustering algorithm can be successfully used directly on the dataset of spectra, and a dimensionality reduction phase with T-SNE is needed beforehand, in which I tested several metrics. Only two of them produced a valid result in a reasonable amount of time: the so-called ‘hamming’ metric and the ‘correlation’ metric. After this dimensionality reduction phase, I run DBSCAN on the reduced dataset to search possible clusters. Using the ’correlation’ metric, I was able to identify one main cluster surrounded by a small set of noise points, which are spectra dissimilar from every other one in the sample. I then applied DBSCAN recursively on the main central cluster, finding a total of 27 clusters organized hierarchically. The objects in each cluster appear to have very similar spectral features, total stellar masses, age, and other physical properties. From these clusters, I also derived a set of 27 spectral templates that could be used to estimate the type of objects and their redshift using cross-correlation techniques.