Data Dimensionality Reduction Using Principal Component Analysis: A Case Study

Srivaths Ramasubramanian, C. N. Sowmyarani, Amogh J Athreya, Anjali Devarajan, Adithya Udaya Shankar, Ramakanth Kumar P · 2024

Training of Machine Learning (ML) models requires huge amounts of data. Usually, in the training of sophisticated models, data-sets can be very computationally intensive on resources. In order to reduce the computational burden of large training data-sets on ML models, the work describes Principal Component Analysis (PCA), a popular dimensionality reduction technique used in ML to reduce the size of training data-sets. This method aims to transform the given set of variables into a new set of components on a lower dimensional subspace. After dimensionality reduction, these components retain the prominent features of data, which leads to reduction in volume and efficient analysis. Inferences on the choice of number of components are presented by analyzing the results obtained from variance contribution plots.

Read the paper · More papers on PaperTik