Similarity measures and distance-based methods
Petr Šmilauer, Jan Lepš · Cambridge University Press eBooks · 2014
In many multivariate methods, one of the first steps is to calculate a matrix of similarities (resemblance measures) either between the cases or between the variables. Although this step is not explicitly done in all the methods, in fact each of the multivariate methods works (even if implicitly) with some similarity measure. The linear ordination methods can be related to several variants of Euclidean distance, while the unimodal (weighted averaging) ordination methods can be related to chi-square distances. The resemblance functions are reviewed in many texts (e.g. Orloci 1978; Gower & Legendre 1986; Ludwig & Reynolds 1988; Legendre & Legendre 2012), so here we will introduce only the most important ones. In this chapter, we will use the following notation: we have n cases (e.g. relevés), containing m response variables (e.g. species). Y ik represents the value (abundance) of the k- th variable ( k = 1, 2,…, m ) in the i- th case ( i= 1, 2,…, n ). In this chapter, we will also refer to (response) variables of the analysed data table as species, to reflect the common nature of data and increase the readability of the text.