Study of dimension reduction methodologies in data mining
Nitika Sharma, Kriti Saroha · 2015
The data mining applications such as bioinformatics, risk management, forensics etc., involves very high dimensional dataset. Due to large number of dimensions, a well known problem of “Curse of Dimensionality” occurs. This problem leads to lower accuracy of machine learning classifiers due to involvement of many insignificant and irrelevant dimensions or features in the dataset. There are many methodologies that are being used to find the Critical Dimensions for a dataset that significantly reduces the number of dimensions. These feature reduction and subset selection methods reduce feature set, that eventually results in high classification accuracy and lower computation cost of machine learning algorithms. This paper surveys the schemes that are majorly used for Dimensionality Reduction mainly focusing Bioinformatics, Agricultural, Gene and Protein Expression datasets. A comparative analysis of surveyed methodologies is also done, based on which, best methodology for a certain type of dataset can be chosen.