Dimensionality Reduction Techniques: Principles, Benefits, and Limitations
Hemanta Kumar Palo, Santanu Kumar Sahoo, Asit Kumar Subudhi · 2021
In machine learning, the effectiveness of the classification and regression analysis is often based on the reliability and dimension of the features representing the patterns. Nevertheless, there are several factors in the feature space which are correlated, overlapped, and can be considered irrelevant. These redundant data may occupy unnecessary storage space, and increase the response time, thus affecting the classification accuracy. The higher the redundant data, the complex it becomes to visualize the training set and simulating on it. Thus, there is a need to reduce the dimension of the data by keeping only relevant information. A 3-D pattern classification task is difficult to contemplate, whereas it is easy to map a 2-D pattern into a simple 2-D space or a 1-D task into a simple line. This is the reason, the dimension reduction techniques have been a growing trend of research in the field of pattern classification and regression. It is a process of reducing the number of features to the desired set by considering the principle variable which can adequately describe the given pattern. This motivates the authors to investigate several dimensional reduction techniques in the light of their principles, advantages, and limitations in this work. It also provides a comparative analysis of the discussed techniques and their suitability on a particular field of application.