Identification of High Leverage Points in High Dimensional Sparse and Non-Sparse Data

Siti Zahariah, Habshah Midi · 2024

The analysis of high dimensional data has become increasingly important in many fields such as applied sciences, engineering, and medicines. For instance, there are tens of thousands of gene expression values available in tumor classification utilizing genomic data; however, the number of arrays is only at the order of ten. Other applications in high dimensional data (HDD) are in chemometrics, fraud detection, climate studies, and satellite processing. HDD that refer to a situation when the number of predictor variables ( p ) is much larger than the sample size ( n ), forms a major statistical challenge in terms of data classification and other statistical analyses. In HDD, a matrix related to some algorithms may become singular. This characteristic leads to a sparsity problem known as the curse of dimensionality phenomena ( Lee et al. 2011 ). In the curse of high dimensionality, conventional statistical methods do not work well, and most of them fail to perform especially when dealing with contaminated data or outliers.

Read the paper · More papers on PaperTik