Random projections fuzzy k-nearest neighbor(RPFKNN) for big data classification

Mihail Popescu, James M. Keller · 2016

As the number of features in pattern recognition applications continuously grows, new algorithms are necessary to reduce the dimensionality of the feature space while producing comparable results. For example, a dynamic area of research, activity recognition, produces large quantities of high-velocity, high-dimensionality data that require real time classification. While dimensionality reduction approaches such as principle component analysis (PCA) and feature selection work well for datasets of reasonable size and dimensionality, they fail on big data. A possible approach to classification of high-dimensionality datasets is to combine a typical classifier, fuzzy k-nearest neighbor in our case (FKNN), with feature reduction by random projection (RP). As opposed to PCA where one projection matrix is computed based on least square optimality, in RP, a projection matrix is chosen at random multiple times. As the random projection procedure is repeated many times, the question is how to aggregate the values of the classifier obtained in each projection. In this paper we present a fusion strategy for RP FKNN, denoted as RPFKNN. The fusion strategy is based on the class membership values produced by FKNN and classification accuracy in each projection. We test RPFKNN on several synthetic and activity recognition datasets.

Read the paper · More papers on PaperTik