A new preprocessing method reduces the dimensionality of classification models

Fatima El Barakaz, Omar Boutkhoum, Abdelmajid El Moutaouakkil · 2019

In data mining classification problems, the higher the number of features, the harder the visualization of the training set is. Working on a dataset with varied factors sounds a big problem for data analyst even after applying classification models, especially when we could not identify the features that have generated such classification. Sometimes, most of these features are correlated, and hence redundant. Dimensionality reduction is the process of reducing the number of random variables under consideration, by obtaining a set of principal variables.

Read the paper · More papers on PaperTik