A Broad Learning based Classification Model for Sparse and High Dimensional Data

Yanqing Ye, Bin Lin, Li Ma, Weilong Yang, Bu Pu, Xiaomin Zhu · 2024

Classifying sparse high-dimensional data is a significant and formidable challenge in the realm of data mining, Traditional machine learning approaches often encounter the issues of dimensionality curse and inefficient computation when processing this type of data. This paper introduces a classification model grounded in the broad learning framework, designed to address the complexities of classifying sparse high-dimensional data. Given the characteristic high dimensionality and sparsity of human gut microbiota data, it offers a distinctive vantage point for diagnosing diseases and monitoring health. Thus, this study applies the broad learning classification model to the data of human gut microbiota. To counteract the sparsity and high dimensionality of microbial data, the data is projected into both the feature and enhancement spaces. By integrating a novel input layer, the internal weight structure can learn through ridge regression approximation and incremental learning that extends into the feature and enhancement spaces. To further enhance the precision of classification, a Broad Learning Classification Model that incorporates both microbial data and meta-features (BLCMs) is proposed. A case study with 21 real-world microbial datasets is conducted to validate the proposed models. The model demonstrates more consistent performance in differentiating patients from healthy individuals when compared to nearest neighbor and random forest classification methods. Furthermore, by incorporating meta-features, BLCMs have achieved notable enhancements in performance. The experimental outcomes indicate that the proposed Broad Learning classification model not only elevates the precision of classification but also substantially decreases computational costs when managing high-dimensional sparse data.

Read the paper · More papers on PaperTik