Analysis of Distance Matrices

Reza Modarres · Statistics & Probability Letters · 2022

Distance matrices and their eigen-structures constitute a fundamental framework for High Dimensional Data Analysis, particularly when the original observations are unavailable or inherently relational. This chapter systematically investigates constant, two-sample, and HDLSS distance matrices, with a focus on outlier detection, imputation of missing distances, change point analysis, and the evaluation of distributional equality. We introduce and analyze two dissimilarity measures, https://www.w3.org/1998/Math/MathML" display="inline"> δ and https://www.w3.org/1998/Math/MathML" display="inline"> ρ , alongside the https://www.w3.org/1998/Math/MathML" display="inline"> L 1 -norm Euclidean distance, deriving their asymptotic eigenvalues for scenarios with one to three groups of observations. Using principal coordinate analysis and Frobenius norm-based statistics, we establish connections between eigenvalues, energy statistics, and tests for shift, scale, and joint shift-scale alternatives. The chapter elucidates the impact of outliers on distance matrices, highlighting potential distortions of eigenvalues and the limitations of existing methods. Moreover, we demonstrate that, in high-dimensional settings, distances within a group converge to a constant structure, enabling rigorous statistical tests for the equality of https://www.w3.org/1998/Math/MathML" display="inline"> K distributions. Applications span bioinformatics, machine learning, cognitive science, and pattern recognition, providing robust tools for dimensionality reduction, clustering, and hypothesis testing. The theoretical developments are complemented with practical guidance for interpreting eigenvalue-based measures and for handling HDLSS data.

Read the paper · More papers on PaperTik