Comment on "Acoustic seabed classification: improved statistical method"

Jon M. Preston, Rodney Lynn Kirlin · Canadian Journal of Fisheries and Aquatic Sciences · 2003

1300 In a discussion of methods for acoustic seabed classification, Legendre et al. (2002) claim to offer improvements over existing techniques and assert that their method “produces statistically better results than the classification method implemented in the QTC [Quester Tangent Corporation] software”. Reasons for this assertion are not given in that paper but are given in an unpublished document. In this paper, we examine the basis for the assertion and discuss whether it should be accepted. The method of Legendre et al. (2002) implements a Kmeans partitioning by an iterative process. They claim this as an advantage over QTC; however, QTC also implements a Kmeans partitioning with an iterative process (see, e.g., Preston et al. 2001), so this cannot be the explanation. Choosing the optimal number of clusters can be problematic, but Legendre et al. report differences between their methods and QTC even when both are computing with the same number of clusters, so the explanation cannot lie entirely here. What is left? The answer, it appears, is that Legendre et al. (2002) wish to minimize the within-group sums of squares using a homogeneous measure of distance to the centre of a cluster (squared Euclidean distances), whereas QTC uses a likelihood-based measure in which the distance to the centre of any cluster is scaled by the variance of that cluster. A one-dimensional example illustrates the point at issue. Suppose we have a population that is an equal mixture of two Gaussian distributions: A, which is distributed as N(0,(0.5)2), and B ~ N(2,(0.1)2). How do we assign an observation at 1.5? Legendre et al. would say that it is a distance of 0.5 from the centre of B and 1.5 from the centre of A and therefore should be assigned to B. QTC would say that is it 5 standard deviations from the centre of B and 3 standard deviations from the centre of A and therefore should be assigned to A. In the example that we have sketched, the statement “x is from A and x = 1.5” has a higher probability density than “x is from B and x = 1.5”; that is, our method has a higher probability of making a correct assignment to a component of the mixture. Therefore, one cannot unambiguously identify Legendre et al.’s “best in the [homogeneous] least-squares sense” with best in the statistical sense. Attempts to prove the latter fail because any statistical test is also based on either homogeneous or variance-based measures; if the clustering method and the test use the same measure, the test scores will often be higher for that reason alone. The paragraphs above summarize our comment. Legendre et al.’s (2002) claim to “statistically better results” is without support, arising as it did from applying a test that used squared Euclidean distances to two clustering methods, one using squared Euclidean distances and one using a likelihood-based measure. What remains is to provide a justification for using non-Euclidean measures for clustering and to extend this example to more dimensions. A reasonable principle for choosing the class assignment is the maximum a posteriori probability (MAP). In other words, a vector x is observed and is to be assigned to one of q classes {ω1,ω2,...,ωq}. MAP selects the class of maximum a posteriori probability from the candidates as

Read the paper · More papers on PaperTik