A manifold learning framework for reducing high-dimensional big text data

Rashed Salem · 2017

The evolution of Big Data and the necessity for analyzing their huge volumes of data raises the fundamental issue of reducing dimensionality. Reducing dimensionality is an important data preparation technique to process large volumes data for data mining, e.g., data classification. The complex representation of high-dimensional data makes the decision space more complicated and timely consuming on building learning models in several applications including Big text classification. This paper considers reducing methods that resolve the high data dimensionality problem, e.g., Isomap, Locally Linear Embedding (LLE) and Singular Vector Decomposition (SVD). It suggests a framework based on manifold learning that merges Random Forest technique, as a learning approach and as feature selection tool, with dimensionality reduction technique to increase the accuracy of predictive analysis for Big text data. Results on Twitter datasets confirm the validity of the suggested learning approach.

Read the paper · More papers on PaperTik